Benchmark Radar
AI BENCHMARK PROFILE

AgenticDataBench

General AIKnowledge & Reasoning

Evaluates LLM-based data agents on realistic data science workflows across 15 domains, with fine-grained ground-truth labels and skill-level scoring.

Released
2026-07-02
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the lack of comprehensive benchmarks for automating data science workflows, providing fine-grained skill-level insights to guide agent development and selection.

Motivation

Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.