Benchmark Radar
AI BENCHMARK PROFILE

ConVBench

General AIMultimodal PerceptionConVBench Team

ConVBench is a vision-centric reasoning benchmark where each image is paired with two logically equivalent questions across six categories (action/state, complex counting, spatial reasoning, causal/intent, commonsense, temporal perception), with metrics for logical consistency and robust accuracy.

Released
2026-07-23
Readiness
Paper only
Primary field
General AI

Why it matters

Reliable visual reasoning requires not just correctness but consistency; this benchmark measures both, offering a more rigorous test for LVLMs and motivating consistency-aware training methods.

Motivation

While Large Vision-Language Models (LVLMs) exhibit strong perceptual capabilities, they remain vulnerable in visual reasoning tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.