AI BENCHMARK PROFILE
StatABench
StatABench evaluates LLMs' statistical analysis capabilities through two components: Stat-Closed, 404 questions across 18 topics in multiple formats, and Stat-Open, 30 complex modeling tasks from competitions, scored via LLM-as-judge.
- Released
- 2026-06-22
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the need for systematic evaluation of LLMs in statistical analysis, revealing performance gaps and challenges in tool-grounded reasoning and modeling.
Motivation
Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.