Benchmark Radar
AI BENCHMARK PROFILE

LADBench

General AIMultimodal PerceptionLADBench Team

LADBench evaluates large vision-language models on detecting logical anomalies in synthetic images across four domains: Residential, Urban, Collaborative, and Nature. It uses a Tiered Prompting Protocol with three progressive disclosure levels and automated scoring.

Released
2026-06-16
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing anomaly benchmarks focus on visual errors, not the physical and social common sense required for open-world deployment. LADBench quantifies how much explicit assistance models need to localize and reason about logical faults, addressing a gap in evaluating sequential multimodal reasoning.

Motivation

Large Vision Language Models (VLMs) excel at visual question answering and semantic grounding, but their capacity for autonomous logical reasoning remains underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.