REAL-Bench
REAL-Bench evaluates vision-driven embodied agents in open-world mobile manipulation across 241 tasks spanning active exploration, visual distraction, articulated manipulation, and interactive disambiguation. The benchmark provides standardized task definitions and a simulator-based evaluation protocol.
- Released
- 2026-07-15
- Readiness
- Runnable
- Primary field
- Robotics & Autonomous Systems
Why it matters
The benchmark targets the gap between simulation and real-world deployment for embodied agents, offering a repeatable evaluation for long-horizon tasks requiring visual grounding and interactive intent disambiguation. It enables systematic comparison of agent frameworks and informs progress toward practical deployment.
Motivation
Real-world deployment of embodied agents requires active exploration, visual grounding, and interactive intent disambiguation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.