AI BENCHMARK PROFILE
DataClawEval
DataClawEval evaluates autonomous data-engineering agents across 100 end-to-end tasks spanning PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL, with deterministic rule-based grading in isolated sandboxes.
- Released
- 2026-07-30
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It provides a reproducible, deterministic evaluation for industrial data-engineering workflows, revealing domain-specific strengths and gaps in agent capabilities.
Motivation
Large language models (LLMs) and LLM-based agents are increasingly being deployed to automate complex workflows, promising to revolutionize data management and processing.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.