Benchmark Radar
AI BENCHMARK PROFILE

MMLU Chat

Finance & EconomicsMathematics & Formal Sciences

Chat-format variant of the Massive Multitask Language Understanding benchmark, evaluating language models across 57 tasks including elementary mathematics, US history, computer science, law, and other professional and academic subjects. This version uses conversational prompting format for model evaluation.

Released
Unknown
Readiness
Paper only
Primary field
Finance & Economics

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.