Benchmark Radar
AI BENCHMARK PROFILE

MeetingToM

General AIMultimodal Perception

MeetingToM evaluates multimodal LLMs on theory-of-mind reasoning in multi-party meetings, including pseudo-consensus detection, across three levels: subject, dyad, and group. It provides a unified evaluation protocol.

Released
2026-07-21
Readiness
Paper only
Primary field
General AI

Why it matters

It covers latent social states and group dynamics often missing in existing ToM benchmarks, revealing limitations in integrating non-verbal cues and inferring hidden attitudes.

Motivation

Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.