ConfidenceBench
A calibration benchmark evaluating verbalized confidence estimates in frontier LLMs using Brier scores across 200 multiple-choice questions in four categories. Scores are elicited via prompting without logits, applicable to closed and open models.
- Released
- 2026-07-10
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The benchmark addresses the need to assess model calibration separately from accuracy, which is critical for trustworthy deployment. The private nature of the questions and lack of public artifacts prevent other teams from running or inspecting the benchmark, limiting its standalone utility.
Motivation
Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.