TriViewBench
TriViewBench evaluates multimodal LLMs on multi-view structural reasoning using synthetic 3D scenes with controlled object count and occlusion. It comprises 1,923 scenes and over 14,000 QA pairs across four complexity levels and three reasoning categories: Local Decision, Object Counting, and Global Recovery.
- Released
- 2026-06-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
TriViewBench provides controlled complexity scaling to isolate structural reasoning capabilities in MLLMs, revealing distinct failure modes and bottlenecks. It enables systematic comparison of models on multi-view spatial reasoning, informing targeted improvements.
Motivation
Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under controlled structural complexity remains poorly understood.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.