Benchmark Radar
AI BENCHMARK PROFILE

DataClawEval

General AIAgentsDataClawEval Team

DataClawEval evaluates autonomous data-engineering agents across 100 end-to-end tasks spanning PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL, with deterministic rule-based grading in isolated sandboxes.

Released
2026-07-30
Readiness
Runnable
Primary field
General AI

Why it matters

It provides a reproducible, deterministic evaluation for industrial data-engineering workflows, revealing domain-specific strengths and gaps in agent capabilities.

Motivation

Large language models (LLMs) and LLM-based agents are increasingly being deployed to automate complex workflows, promising to revolutionize data management and processing.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.