Benchmark Radar
AI BENCHMARK PROFILE

MSQA

General AIKnowledge & Reasoning

MSQA evaluates 1,064 natively sourced questions across 11 language groups, five cultural dimensions, and three difficulty tiers, testing cultural knowledge in a multilingual context.

Released
2026-07-01
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses the gap in measuring cultural alignment separately from language ability, revealing that models often fail to achieve cultural competence despite multilingual fluency.

Motivation

Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.