Benchmark Radar
AI BENCHMARK PROFILE

Blind-Spots-Bench

General AIMultimodal Perception

Evaluates reasoning blind spots in language, vision-language, and image-generation models across 235 samples with structured reference solutions and taxonomy.

Released
2026-07-09
Readiness
Runnable
Primary field
General AI

Why it matters

Serves as a diagnostic stress test exposing tasks that humans find easy but AI models struggle with, highlighting gaps not captured by existing benchmarks.

Motivation

Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.