Benchmark Radar
AI BENCHMARK PROFILE

PIVOTSBench

General AIMultimodal Perception

PIVOTSBench evaluates multimodal large language models on fine-grained interpersonal relationship reasoning. It includes tasks predicting bidirectional relationship dimensions from videos and auxiliary tasks on visual cue identification.

Released
2026-06-22
Readiness
Inspectable
Primary field
General AI

Why it matters

This benchmark addresses the lack of evaluation for multimodal social reasoning, providing a standardized test for model capabilities in understanding nuanced interpersonal cues, which is crucial for developing AI that interacts naturally in social contexts.

Motivation

Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.