LongDS-Bench
LongDS-Bench evaluates long-horizon, multi-turn data analysis tasks where agents must maintain, update, restore, and compose evolving analytical states. It comprises 68 tasks from real-world Kaggle notebooks spanning 2,225 turns across six domains, with an average dependency span of 11.3 turns.
- Released
- 2026-05-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks focus on isolated or short interactive tasks, leaving long-horizon analytical state management untested. This benchmark reveals a critical bottleneck in agent performance, where errors concentrate in later turns and additional interaction steps do not reliably improve accuracy, aiding development of more reliable agentic systems.
Motivation
Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.