Benchmark Radar
AI BENCHMARK PROFILE

Avalon-ToM-Bench

General AIKnowledge & Reasoning

Avalon-ToM-Bench evaluates fine-grained theory of mind in LLMs using a 2x2 taxonomy of epistemic/motivational reasoning crossed with inference/action, via human-crafted perspective-constrained queries from The Resistance: Avalon.

Released
2026-08-10
Readiness
Paper only
Primary field
General AI

Why it matters

It provides a diagnostic decomposition of ToM abilities, distinguishing reasoning, expression, and policy, and offers insights into training and inference interventions for improving social reasoning.

Motivation

Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.