ReactBench
ReactBench evaluates multimodal large language models on cause-driven hallucination through four targeted tasks (Relational Erasure, Counterfactual Attribute, Alteration Tracing, Dense Counting) with exam-style evaluation and chain-of-thought reasoning for sub-cause identification.
- Released
- 2026-05-28
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing hallucination benchmarks measure outcomes rather than causes. ReactBench provides a systematic testbed to diagnose specific failure modes like co-occurrence bias and fine-grained perceptual bottlenecks, offering interpretable insights for model robustness.
Motivation
While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinations, producing responses that are inconsistent with the visual input.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.