Benchmark Radar
AI BENCHMARK PROFILE

EarthVerse

General AIKnowledge & Reasoning

Evaluates scientific agents on 405 reproducible tasks across 199 documented natural hazard events, scoring fine-grained answer units and task-specific rubrics.

Released
2026-08-24
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a rigorous, provenance-focused measure of end-to-end scientific reliability for agents dealing with dynamic Earth systems and natural hazards.

Motivation

Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.