DataGovBench
DataGovBench evaluates LLMs on real-world data analysis using government open data. It comprises Table QA (complex decomposable questions with textual or visual answers) and Table Insight (exploratory data analysis producing expert-level findings). Scoring uses unambiguous reference-based metrics.
- Released
- 2026-07-07
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing LLM benchmarks miss real-world data complexities like large multi-table datasets and exploratory insight discovery. DataGovBench provides a challenging evaluation to gauge practical readiness of LLMs for data analytics tasks.
Motivation
Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.