Benchmark Radar
AI BENCHMARK PROFILE

LongDS-Bench

General AIKnowledge & ReasoningZhejiang University NLP Lab (ZJUNLP)

LongDS-Bench evaluates long-horizon, multi-turn data analysis tasks where agents must maintain, update, restore, and compose evolving analytical states. It comprises 68 tasks from real-world Kaggle notebooks spanning 2,225 turns across six domains, with an average dependency span of 11.3 turns.

Released
2026-05-28
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks focus on isolated or short interactive tasks, leaving long-horizon analytical state management untested. This benchmark reveals a critical bottleneck in agent performance, where errors concentrate in later turns and additional interaction steps do not reliably improve accuracy, aiding development of more reliable agentic systems.

Motivation

Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.