CDR-Bench
CDR-Bench is a benchmark of 3,462 tasks for evaluating large language models on faithful execution of compositional, order-sensitive data refinement recipes across four domains and 29 operators, with deterministic reference outputs for exact evaluation.
- Released
- 2026-06-30
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks leave unclear whether LLMs can directly execute multi-step, order-sensitive data refinement tasks; this benchmark provides a reusable, deterministic evaluation protocol to assess procedural faithfulness in compositional text processing.
Motivation
Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing operators determine the outcome.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.