Benchmark Radar
AI BENCHMARK PROFILE

DisasterBench

General AIKnowledge & ReasoningDisasterBench Team

DisasterBench is a benchmark for evaluating structured multi-agent planning over disaster-response tools. It includes 233 expert-verified tasks, 26 agents, 81 typed compatibility edges, and 5 planning paradigms. It uses First-Point-of-Failure (FPoF) for step-level failure attribution.

Released
2026-05-27
Readiness
Runnable
Primary field
General AI

Why it matters

Disaster response requires orchestrating heterogeneous AI tools into executable workflows. DisasterBench tests grounding under typed interface constraints, highlighting gaps between semantic reasoning and execution consistency, and providing diagnostics for failure attribution.

Motivation

Disasters cause severe societal impacts, demanding rapid coordination of heterogeneous AI tools, from satellite analysis to flood prediction and damage assessment, into coherent multi-step workflows.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.