UrduMMLU
UrduMMLU evaluates Urdu language understanding through 26,431 multiple-choice questions across 26 subjects and five domains, sourced from native educational materials. Accuracy under zero-shot and few-shot prompting protocols serves as the primary scoring metric.
- Released
- 2026-06-05
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Urdu, spoken by over 230 million people, lacks broad MMLU-style evaluation from native sources. UrduMMLU addresses this gap by testing models on region-specific knowledge, revealing uneven performance across subjects and providing a public benchmark for comparing LLMs on Urdu understanding.
Motivation
Meaningful multilingual evaluation must test models in the target language and educational context.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.