DSAgentBench
DSAgentBench evaluates language agents on end-to-end data-science workflows in real computer environments. It comprises 275 tasks spanning data wrangling, exploration, modeling, visualization, and validation, with deterministic evaluators verifying analytical correctness, visual outputs, and model performance.
- Released
- 2026-08-11
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks lack real-computer interaction and fail to capture the multi-stage, multi-tool nature of data-science practice. DSAgentBench provides a realistic environment for assessing whether agents can automate complete workflows, highlighting a significant capability gap and offering a foundation for developing grounded, verifiable autonomous agents.
Motivation
Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.