WorldCoder-Bench
Evaluates autonomous physically grounded 3D world synthesis from natural language, with 2,026 expert-curated tasks across Simulation, Rendering, and Application scenarios, using execution-based verification via StateProbe to check hidden behavioral contracts over runtime states.
- Released
- 2026-06-01
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing web-generation benchmarks observe only pixels or DOM nodes, missing the mechanics of Three.js worlds inside canvas elements. This benchmark provides a protocol for verifying hidden contracts, enabling assessment of correctness-adjusted cost and time efficiency for 3D world synthesis.
Motivation
Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural language.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.