Benchmark Radar
AI BENCHMARK PROFILE

CODA-BENCH

General AIAgentsRUC-DataLab

CODA-Bench is a benchmark for evaluating AI agents on data-intensive analytical tasks in a Linux sandbox with 1,009 tasks across 31 communities, requiring data discovery, code generation, and correct answers.

Released
2026-06-13
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks evaluate code or data capabilities in isolation. CODA-Bench jointly evaluates both, reflecting real development scenarios with complex file systems and large-scale data, revealing gaps in agentic data intelligence.

Motivation

Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that capture the complexity of real-world development.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.