scBench-Long
scBench-Long evaluates long-horizon single-cell biology reasoning. Agents must recover scientific conclusions from raw or near-raw data without prescribed methods. It contains 21 evaluations spanning diverse biological contexts, with deterministic grading and trajectory rubrics.
- Released
- 2026-06-25
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Existing AI-biology benchmarks measure broad knowledge or local steps; scBench-Long assesses end-to-end scientific claim production. It provides a reusable evaluation for long-horizon reasoning in single-cell data analysis.
Motivation
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.