AI BENCHMARK PROFILE
PIVOTSBench
PIVOTSBench evaluates multimodal large language models on fine-grained interpersonal relationship reasoning. It includes tasks predicting bidirectional relationship dimensions from videos and auxiliary tasks on visual cue identification.
- Released
- 2026-06-22
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
This benchmark addresses the lack of evaluation for multimodal social reasoning, providing a standardized test for model capabilities in understanding nuanced interpersonal cues, which is crucial for developing AI that interacts naturally in social contexts.
Motivation
Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.