Benchmark Radar
AI BENCHMARK PROFILE

AutoWorldModel-Bench

General AIKnowledge & ReasoningElectronic Arts

AutoWorldModel-Bench evaluates coding agents on autonomously improving a base world model under fixed compute budget across eight game environments. Uses structured-state representation, with held-out test split and closed-loop iteration.

Released
2026-07-20
Readiness
Inspectable
Primary field
General AI

Why it matters

Addresses the gap in agent benchmarks for open-ended research tasks, enabling comparison of agents on research-like workflows rather than engineering-to-spec tasks.

Motivation

World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.