LADBench
LADBench evaluates large vision-language models on detecting logical anomalies in synthetic images across four domains: Residential, Urban, Collaborative, and Nature. It uses a Tiered Prompting Protocol with three progressive disclosure levels and automated scoring.
- Released
- 2026-06-16
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing anomaly benchmarks focus on visual errors, not the physical and social common sense required for open-world deployment. LADBench quantifies how much explicit assistance models need to localize and reason about logical faults, addressing a gap in evaluating sequential multimodal reasoning.
Motivation
Large Vision Language Models (VLMs) excel at visual question answering and semantic grounding, but their capacity for autonomous logical reasoning remains underexplored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.