Benchmark Radar
AI BENCHMARK PROFILE

OmniToM

General AIKnowledge & Reasoning

OmniToM evaluates theory of mind in LLMs by requiring explicit belief modeling, extracting belief propositions and labeling them with seven-dimensional schema labels across 895 stories.

Released
2026-05-25
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses the gap of end-point question answering in ToM evaluation by forcing explicit mental-state representation, potentially revealing actor-specific belief-tracking bottlenecks.

Motivation

Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-point question answering, where performance is judged solely by the final answer to a social reasoning query.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.