Benchmark Radar
AI BENCHMARK PROFILE

AgentWorldBench

General AIKnowledge & Reasoning

AgentWorldBench evaluates language world models on simulation fidelity across 7 domains (MCP, Search, Terminal, SWE, Android, Web, OS) using real-world trajectories and rubric-based scoring.

Released
2026-06-23
Readiness
Runnable
Primary field
General AI

Why it matters

This benchmark addresses the lack of systematic evaluation for language world models in agentic environments, providing a standardized protocol that enables direct comparison and guides development of general agents.

Motivation

A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.