Benchmark Radar
AI BENCHMARK PROFILE

UrduMMLU

General AIKnowledge & Reasoning

UrduMMLU evaluates Urdu language understanding through 26,431 multiple-choice questions across 26 subjects and five domains, sourced from native educational materials. Accuracy under zero-shot and few-shot prompting protocols serves as the primary scoring metric.

Released
2026-06-05
Readiness
Paper only
Primary field
General AI

Why it matters

Urdu, spoken by over 230 million people, lacks broad MMLU-style evaluation from native sources. UrduMMLU addresses this gap by testing models on region-specific knowledge, revealing uneven performance across subjects and providing a public benchmark for comparing LLMs on Urdu understanding.

Motivation

Meaningful multilingual evaluation must test models in the target language and educational context.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.