Benchmark Radar
AI BENCHMARK PROFILE

DSAgentBench

General AIKnowledge & Reasoningvis-nlp

DSAgentBench evaluates language agents on end-to-end data-science workflows in real computer environments. It comprises 275 tasks spanning data wrangling, exploration, modeling, visualization, and validation, with deterministic evaluators verifying analytical correctness, visual outputs, and model performance.

Released
2026-08-11
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks lack real-computer interaction and fail to capture the multi-stage, multi-tool nature of data-science practice. DSAgentBench provides a realistic environment for assessing whether agents can automate complete workflows, highlighting a significant capability gap and offering a foundation for developing grounded, verifiable autonomous agents.

Motivation

Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.