Benchmark Radar
AI BENCHMARK PROFILE

LongEarth-Bench

General AIMultimodal Perception

LongEarth-Bench evaluates vision-language models on long-horizon Earth observation reasoning, with ~120k QA samples from 117k images, sequences avg 15.14 frames (up to 30), covering 12 tasks in evolution summarization, spatial reasoning, anomaly identification, and logical prediction.

Released
2026-08-13
Readiness
Paper only
Primary field
General AI

Why it matters

LongEarth-Bench addresses the lack of benchmarks for long-sequence Earth observation reasoning, providing a more realistic and challenging evaluation for models designed to analyze multi-stage geographic changes.

Motivation

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.