Benchmark Radar
AI BENCHMARK PROFILE

mmPISA-bench

General AIKnowledge & Reasoning

mmPISA-bench consists of 25 multiple-choice questions from PISA in 43 languages with official and machine translations, used to evaluate LLMs' reasoning across languages.

Released
2026-06-05
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses multilingual reasoning evaluation but the small scale and focus on proprietary models limit its utility as a general benchmark.

Motivation

We introduce mmPISA-bench, a compact high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment (PISA).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.