Benchmark Radar
AI BENCHMARK PROFILE

StatABench

General AIKnowledge & Reasoning

StatABench evaluates LLMs' statistical analysis capabilities through two components: Stat-Closed, 404 questions across 18 topics in multiple formats, and Stat-Open, 30 complex modeling tasks from competitions, scored via LLM-as-judge.

Released
2026-06-22
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the need for systematic evaluation of LLMs in statistical analysis, revealing performance gaps and challenges in tool-grounded reasoning and modeling.

Motivation

Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.