ERQA-Plus
ERQA-Plus evaluates embodied reasoning in AI systems with 1,766 question-answer instances grounded in 711 robot-centric images, organized by a taxonomy covering perceptual, action-centric, social-interaction, navigation-environmental, and commonsense reasoning. Scoring uses overall accuracy and SBERT similarity.
- Released
- 2026-06-16
- Readiness
- Runnable
- Primary field
- Robotics & Autonomous Systems
Why it matters
Existing visual and embodied QA benchmarks often lack control over reasoning dependencies, making it hard to distinguish genuine embodied reasoning from shortcut-driven pattern matching. ERQA-Plus provides a fine-grained diagnostic to identify strengths and weaknesses in specific reasoning categories, informing model development and deployment decisions.
Motivation
Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visual observations.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.