Benchmark Radar
AI BENCHMARK PROFILE

CITBench

General AIKnowledge & Reasoning

CITBench evaluates LLMs on interactive tabular data processing, covering table matching, cleaning, augmentation, and transformation across 18 task types and 1,296 instances.

Released
2026-06-29
Readiness
Paper only
Primary field
General AI

Why it matters

Tabular data processing benchmarks often focus on single-turn reasoning; interactive multi-turn settings remain underevaluated.

Motivation

Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.