RescueBench
RescueBench evaluates embodied search-and-rescue agents in simulated photo-realistic environments. It comprises a four-stage pipeline: multimodal exploration, target rescue, memory-guided return, and final handoff. Five difficulty levels vary environmental complexity, clue ambiguity, and spatial hierarchy. Automatic episode generation and annotation support scalable evaluation. A unified benchmark framework and runner scripts provide standardized scoring.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- Robotics & Autonomous Systems
Why it matters
Search-and-rescue benchmarks typically test capabilities in isolation. RescueBench addresses the gap of composite workflows where failures may compound across stages. It provides stage-level diagnostics to identify bottlenecks (e.g., exploration, memory) separately from end-to-end performance, helping practitioners target improvements in embodied agents for realistic rescue tasks.
Motivation
Search-and-rescue (SAR) requires embodied agents to explore unfamiliar environments under multimodal uncertainty, perform multi-stage interactions, and retrieve spatial memory over long horizons.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.