AI BENCHMARK PROFILE
Blind-Spots-Bench
Evaluates reasoning blind spots in language, vision-language, and image-generation models across 235 samples with structured reference solutions and taxonomy.
- Released
- 2026-07-09
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Serves as a diagnostic stress test exposing tasks that humans find easy but AI models struggle with, highlighting gaps not captured by existing benchmarks.
Motivation
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.