AI BENCHMARK PROFILE
AutoWorldModel-Bench
AutoWorldModel-Bench evaluates coding agents on autonomously improving a base world model under fixed compute budget across eight game environments. Uses structured-state representation, with held-out test split and closed-loop iteration.
- Released
- 2026-07-20
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Addresses the gap in agent benchmarks for open-ended research tasks, enabling comparison of agents on research-like workflows rather than engineering-to-spec tasks.
Motivation
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.