Benchmark Radar
AI BENCHMARK PROFILE

SpatialWorld

General AIAgentsSpatialWorld Team

SpatialWorld evaluates multimodal agents on interactive spatial reasoning in 760 real-world tasks across eight simulation backends. Agents operate under vision-only partial observability, using a unified text-based action interface. Performance is measured via terminal-state verifiers for task success rate and step efficiency.

Released
2026-06-08
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks rely on passive VQA or simulator-specific pipelines, failing to assess interactive spatial understanding. SpatialWorld provides a unified, simulator-agnostic protocol with human-validated evaluation, revealing that current models achieve low success rates, highlighting gaps in active exploration and long-horizon planning.

Motivation

Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.