Benchmark Radar
AI BENCHMARK PROFILE

WorldCoder-Bench

General AIMultimodal PerceptionWorldCoder-Bench team

Evaluates autonomous physically grounded 3D world synthesis from natural language, with 2,026 expert-curated tasks across Simulation, Rendering, and Application scenarios, using execution-based verification via StateProbe to check hidden behavioral contracts over runtime states.

Released
2026-06-01
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing web-generation benchmarks observe only pixels or DOM nodes, missing the mechanics of Three.js worlds inside canvas elements. This benchmark provides a protocol for verifying hidden contracts, enabling assessment of correctness-adjusted cost and time efficiency for 3D world synthesis.

Motivation

Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural language.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.