AI BENCHMARK PROFILE
AgentWorldBench
AgentWorldBench evaluates language world models on simulation fidelity across 7 domains (MCP, Search, Terminal, SWE, Android, Web, OS) using real-world trajectories and rubric-based scoring.
- Released
- 2026-06-23
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
This benchmark addresses the lack of systematic evaluation for language world models in agentic environments, providing a standardized protocol that enables direct comparison and guides development of general agents.
Motivation
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.