SSMNBench
SSMNBench is a diagnostic benchmark for cross-view human and human-object understanding, comprising 3,300 QA pairs categorized into Single-View Sufficiency (SVS) and Multi-View Necessity (MVN) tasks. It evaluates multimodal LLMs by perturbing view availability to assess distraction robustness and cross-view evidence integration.
- Released
- 2026-06-24
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Addresses the gap in evaluating genuine cross-view synthesis versus reliance on single-image semantics in MLLMs. It provides a rigorous framework for diagnosing limitations in cross-view understanding, guiding development of multimodal architectures for complex scenes.
Motivation
Multimodal Large Language Models (MLLMs) have shown remarkable progress in single-image perception, yet their ability to reason about complex cross-view human-centric scenes remains largely unverified.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.