Benchmark Radar
AI BENCHMARK PROFILE

Obshazard-bench

General AIMultimodal Perception

Obshazard-bench is a real-time benchmark for disaster intelligence, integrating raw satellite and ground-station data. It covers 8 disaster categories and 28 sub-categories across 60+ countries, with VQA samples and a three-stage evaluation taxonomy: Predictive Crisis Anticipation, Active Evolution Reasoning, and Multi-faceted Impact Quantification.

Released
2026-06-24
Readiness
Paper only
Primary field
General AI

Why it matters

Fills the gap in evaluating MLLMs for operational disaster response, which require real-time reasoning from raw observation streams. It provides a realistic testbed for assessing decision-support capabilities in evolving emergencies.

Motivation

Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.