AI BENCHMARK PROFILE
ComplexityWorld
ComplexityWorld is a benchmark of 390 visual decision-making tasks across 39 worlds, scored by an executable verifier. It evaluates VLMs on tasks requiring global constraint satisfaction.
- Released
- 2026-08-05
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It targets a persistent visual-to-decision bottleneck in VLMs, but the lack of public artifacts and scoring details hinders independent verification.
Motivation
Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.