Obshazard-bench
Obshazard-bench is a real-time benchmark for disaster intelligence, integrating raw satellite and ground-station data. It covers 8 disaster categories and 28 sub-categories across 60+ countries, with VQA samples and a three-stage evaluation taxonomy: Predictive Crisis Anticipation, Active Evolution Reasoning, and Multi-faceted Impact Quantification.
- Released
- 2026-06-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Fills the gap in evaluating MLLMs for operational disaster response, which require real-time reasoning from raw observation streams. It provides a realistic testbed for assessing decision-support capabilities in evolving emergencies.
Motivation
Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.