Benchmark Radar
AI BENCHMARK PROFILE

RRS-10K

General AIMultimodal Perception

RRS-10K is a benchmark for rare remote sensing image interpretation containing 10,738 military-related images with multiple format question-answer pairs, organized into three capability dimensions and 20 leaf tasks covering perception, reasoning, and robustness.

Released
2026-07-13
Readiness
Paper only
Primary field
General AI

Why it matters

Current remote sensing benchmarks are dominated by common scenes, limiting understanding of VLM performance on rare, long-tail scenarios. RRS-10K provides a standardized evaluation for this gap, enabling systematic analysis of failure modes and guiding development of more reliable remote sensing VLMs.

Motivation

Vision-language models (VLMs) have achieved strong performance on general remote sensing tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.