Benchmark Radar
AI BENCHMARK PROFILE

DataGovBench

General AIKnowledge & ReasoningDataGovBench Team

DataGovBench evaluates LLMs on real-world data analysis using government open data. It comprises Table QA (complex decomposable questions with textual or visual answers) and Table Insight (exploratory data analysis producing expert-level findings). Scoring uses unambiguous reference-based metrics.

Released
2026-07-07
Readiness
Runnable
Primary field
General AI

Why it matters

Existing LLM benchmarks miss real-world data complexities like large multi-table datasets and exploratory insight discovery. DataGovBench provides a challenging evaluation to gauge practical readiness of LLMs for data analytics tasks.

Motivation

Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.