Benchmark Radar
AI BENCHMARK PROFILE

ComplexityWorld

General AIMultimodal Perception

ComplexityWorld is a benchmark of 390 visual decision-making tasks across 39 worlds, scored by an executable verifier. It evaluates VLMs on tasks requiring global constraint satisfaction.

Released
2026-08-05
Readiness
Paper only
Primary field
General AI

Why it matters

It targets a persistent visual-to-decision bottleneck in VLMs, but the lack of public artifacts and scoring details hinders independent verification.

Motivation

Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.