Benchmark Radar
AI BENCHMARK PROFILE

SpatialBench-Long

General AIKnowledge & Reasoning

SpatialBench-Long evaluates AI agents on long-horizon spatial biology tasks, requiring recovery of biological claims from raw data across 24 evaluations.

Released
2026-05-27
Readiness
Paper only
Primary field
General AI

Why it matters

Tests agents' ability to synthesize scientific conclusions from complex spatial data, but no public data or code release is mentioned.

Motivation

AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or localized analysis steps rather than end-to-end scientific reasoning over spatial measurements.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.