ERGeoBench
ERGeoBench evaluates vision-driven embodied geo-localization in MLLMs with 2,207 globally distributed street-view panoramas under single-view, panorama-view, and embodied-view settings. It measures foundational perception, spatial awareness, common sense reasoning, and geo-localization reasoning.
- Released
- 2026-05-29
- Readiness
- Inspectable
- Primary field
- Robotics & Autonomous Systems
Why it matters
Embodied geo-localization is underexplored due to lack of fine-grained evaluation. ERGeoBench provides a unified diagnostic framework that reveals current MLLMs struggle with fine-grained perceptual operations and metric localization, supporting progress in integrated perception and spatial reasoning.
Motivation
Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.