Benchmark Radar
AI BENCHMARK PROFILE

scBench-Long

Health & Life SciencesKnowledge & Reasoning

scBench-Long evaluates long-horizon single-cell biology reasoning. Agents must recover scientific conclusions from raw or near-raw data without prescribed methods. It contains 21 evaluations spanning diverse biological contexts, with deterministic grading and trajectory rubrics.

Released
2026-06-25
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Existing AI-biology benchmarks measure broad knowledge or local steps; scBench-Long assesses end-to-end scientific claim production. It provides a reusable evaluation for long-horizon reasoning in single-cell data analysis.

Motivation

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.