SoMBench
SoMBench evaluates social intelligence in large language models across 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms, with 284 shared scenarios and 3,481 expert-verified instances.
- Released
- 2026-07-26
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
SoMBench targets the gap in evaluating LLMs' social intelligence, providing a structured benchmark to measure capabilities that are increasingly important for long-term deployment in human environments.
Motivation
As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.