Benchmark Radar
AI BENCHMARK PROFILE

GroupToM-Bench

General AIMultimodal Perception

GroupToM-Bench is a multimodal benchmark for evaluating group-level theory of mind through a seven-level cognitive audit framework.

Released
2026-06-02
Readiness
Paper only
Primary field
General AI

Why it matters

It probes a gap in social cognition capabilities of multimodal LLMs, with implications for understanding emergent group behavior.

Motivation

True general intelligence requires not only a model of the physical world but also a social world model: the capacity to infer how individual mental states interact and crystallize into group-level outcomes.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.