Benchmark Radar
AI BENCHMARK PROFILE

RescueBench

Robotics & Autonomous SystemsRobotics & Embodied Intelligence

RescueBench evaluates embodied search-and-rescue agents in simulated photo-realistic environments. It comprises a four-stage pipeline: multimodal exploration, target rescue, memory-guided return, and final handoff. Five difficulty levels vary environmental complexity, clue ambiguity, and spatial hierarchy. Automatic episode generation and annotation support scalable evaluation. A unified benchmark framework and runner scripts provide standardized scoring.

Released
2026-06-01
Readiness
Runnable
Primary field
Robotics & Autonomous Systems

Why it matters

Search-and-rescue benchmarks typically test capabilities in isolation. RescueBench addresses the gap of composite workflows where failures may compound across stages. It provides stage-level diagnostics to identify bottlenecks (e.g., exploration, memory) separately from end-to-end performance, helping practitioners target improvements in embodied agents for realistic rescue tasks.

Motivation

Search-and-rescue (SAR) requires embodied agents to explore unfamiliar environments under multimodal uncertainty, perform multi-stage interactions, and retrieve spatial memory over long horizons.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.