Benchmark Radar
AI BENCHMARK PROFILE

SoMBench

General AIKnowledge & Reasoning

SoMBench evaluates social intelligence in large language models across 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms, with 284 shared scenarios and 3,481 expert-verified instances.

Released
2026-07-26
Readiness
Paper only
Primary field
General AI

Why it matters

SoMBench targets the gap in evaluating LLMs' social intelligence, providing a structured benchmark to measure capabilities that are increasingly important for long-term deployment in human environments.

Motivation

As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.