Benchmark Radar
AI BENCHMARK PROFILE

TriViewBench

General AIMultimodal Perception

TriViewBench evaluates multimodal LLMs on multi-view structural reasoning using synthetic 3D scenes with controlled object count and occlusion. It comprises 1,923 scenes and over 14,000 QA pairs across four complexity levels and three reasoning categories: Local Decision, Object Counting, and Global Recovery.

Released
2026-06-24
Readiness
Paper only
Primary field
General AI

Why it matters

TriViewBench provides controlled complexity scaling to isolate structural reasoning capabilities in MLLMs, revealing distinct failure modes and bottlenecks. It enables systematic comparison of models on multi-view spatial reasoning, informing targeted improvements.

Motivation

Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under controlled structural complexity remains poorly understood.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.