Benchmark Radar
AI BENCHMARK PROFILE

MMLU-redux-2.0

General AIMathematics & Formal Sciences

A curated version of the MMLU benchmark featuring manually re-annotated 5,700 questions across 57 subjects to identify and correct errors in the original dataset. Addresses the 6.49% error rate found in MMLU and provides more reliable evaluation metrics for language models.

Released
Unknown
Readiness
Paper only
Primary field
General AI

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.