Benchmark Radar
AI BENCHMARK PROFILE

NumerosityVLM

General AIMultimodal Perceptionfuy3

Evaluates zero-shot numerosity perception in vision-language models on 10,800 synthetic images across six controlled conditions manipulating object size, spatial arrangement, and numerosity while ablating texture, shape, and color.

Released
2026-08-15
Readiness
Runnable
Primary field
General AI

Why it matters

Isolates numerosity from correlated visual features, enabling targeted diagnosis of number understanding in VLMs and revealing that architecture-driven language components limit performance.

Motivation

Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.