Benchmark Radar
AI BENCHMARK PROFILE

ReactBench

General AIMultimodal PerceptionReactBench Project

ReactBench evaluates multimodal large language models on cause-driven hallucination through four targeted tasks (Relational Erasure, Counterfactual Attribute, Alteration Tracing, Dense Counting) with exam-style evaluation and chain-of-thought reasoning for sub-cause identification.

Released
2026-05-28
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing hallucination benchmarks measure outcomes rather than causes. ReactBench provides a systematic testbed to diagnose specific failure modes like co-occurrence bias and fine-grained perceptual bottlenecks, offering interpretable insights for model robustness.

Motivation

While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinations, producing responses that are inconsistent with the visual input.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.