AI BENCHMARK PROFILE
MSQA
MSQA evaluates 1,064 natively sourced questions across 11 language groups, five cultural dimensions, and three difficulty tiers, testing cultural knowledge in a multilingual context.
- Released
- 2026-07-01
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It addresses the gap in measuring cultural alignment separately from language ability, revealing that models often fail to achieve cultural competence despite multilingual fluency.
Motivation
Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.