Benchmark Radar
AI BENCHMARK PROFILE

OpenAI MMLU

Finance & EconomicsMathematics & Formal Sciences

MMLU (Massive Multitask Language Understanding) is a comprehensive benchmark that measures a text model's multitask accuracy across 57 diverse academic and professional subjects. The test covers elementary mathematics, US history, computer science, law, morality, business ethics, clinical knowledge, and many other domains spanning STEM, humanities, social sciences, and professional fields. To attain high accuracy, models must possess extensive world knowledge and problem-solving ability.

Released
Unknown
Readiness
Paper only
Primary field
Finance & Economics

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.