AI BENCHMARK PROFILE
CITBench
CITBench evaluates LLMs on interactive tabular data processing, covering table matching, cleaning, augmentation, and transformation across 18 task types and 1,296 instances.
- Released
- 2026-06-29
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Tabular data processing benchmarks often focus on single-turn reasoning; interactive multi-turn settings remain underevaluated.
Motivation
Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.