Embodied3DBench
Embodied3DBench is a robot-centric benchmark for low-level spatial intelligence in embodied 3D environments. It includes 6 task categories: Grounding, Spatial Relation Prediction, Multi-view Correspondence, Affordance Prediction, Grasp Point Prediction, and Trajectory Prediction, with 21k QA pairs across 12 subcategories.
- Released
- 2026-05-27
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Current VLMs show strong high-level spatial reasoning but lack interaction-oriented perception. Embodied3DBench reveals this gap and provides a scalable training dataset for improvement, enabling systematic evaluation and advancement of interaction-aware multimodal systems.
Motivation
Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments?
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.