Benchmark Radar
AI BENCHMARK PROFILE

ConfidenceBench

General AIKnowledge & Reasoning

A calibration benchmark evaluating verbalized confidence estimates in frontier LLMs using Brier scores across 200 multiple-choice questions in four categories. Scores are elicited via prompting without logits, applicable to closed and open models.

Released
2026-07-10
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses the need to assess model calibration separately from accuracy, which is critical for trustworthy deployment. The private nature of the questions and lack of public artifacts prevent other teams from running or inspecting the benchmark, limiting its standalone utility.

Motivation

Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.