AI BENCHMARK PROFILE
SpatialBench-Long
SpatialBench-Long evaluates AI agents on long-horizon spatial biology tasks, requiring recovery of biological claims from raw data across 24 evaluations.
- Released
- 2026-05-27
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Tests agents' ability to synthesize scientific conclusions from complex spatial data, but no public data or code release is mentioned.
Motivation
AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or localized analysis steps rather than end-to-end scientific reasoning over spatial measurements.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.