LH-AVLN
LH-AVLN is a benchmark for long-horizon navigation that combines multi-goal missions, heterogeneous goal specifications, and persistent spatialized audio cues. Agents must execute missions of two to four goals specified by category, language, or reference image, using RGB-D, pose, and binaural audio in indoor 3D environments. It supports ordered and unordered missions with alternating goal-associated sounds.
- Released
- 2026-07-04
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Existing navigation benchmarks do not integrate long-horizon missions with audio-visual cues and heterogeneous goal types, while audio-visual tasks are typically single-goal. LH-AVLN addresses this gap by providing a multi-goal environment with acoustic guidance and distractors, enabling more realistic evaluation of embodied agents in complex missions.
Motivation
Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus on a single goal.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.