Benchmark Radar
AI BENCHMARK PROFILE

ChronoBench

General AIMultimodal PerceptionIntelliSensing Lab

ChronoBench evaluates long-term temporal understanding in remote sensing across four progressive cognitive levels: land cover perception, temporal recognition, long-term memory, and spatio-temporal reasoning. It comprises 12 sub-tasks and 17,689 QA pairs over 3,469 images spanning 500 regions across 39 U.S. cities.

Released
2026-07-17
Readiness
Runnable
Primary field
General AI

Why it matters

Existing remote sensing benchmarks typically focus on static or bi-temporal analysis, lacking a systematic dissection of long-term temporal competencies. ChronoBench provides a multidimensional evaluation that isolates specific cognitive bottlenecks, enabling targeted model improvement. Its integration with lmms-eval supports standardized comparison of multimodal LLMs for satellite image time series.

Motivation

Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.