DisasterBench
DisasterBench is a benchmark for evaluating structured multi-agent planning over disaster-response tools. It includes 233 expert-verified tasks, 26 agents, 81 typed compatibility edges, and 5 planning paradigms. It uses First-Point-of-Failure (FPoF) for step-level failure attribution.
- Released
- 2026-05-27
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Disaster response requires orchestrating heterogeneous AI tools into executable workflows. DisasterBench tests grounding under typed interface constraints, highlighting gaps between semantic reasoning and execution consistency, and providing diagnostics for failure attribution.
Motivation
Disasters cause severe societal impacts, demanding rapid coordination of heterogeneous AI tools, from satellite analysis to flood prediction and damage assessment, into coherent multi-step workflows.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.