AI BENCHMARK PROFILE
NOLLI
NOLLI is a procedurally generated English-Korean puzzle benchmark with 15 puzzle types (25 tasks, 7,500 items). Each instance is seed-regenerable, verified to have a unique solution, and scored deterministically. Difficulty is calibrated behaviorally to target accuracy bands.
- Released
- 2026-08-05
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
NOLLI addresses the lack of controlled cross-lingual benchmarks that separate presentation language from reasoning difficulty, enabling diagnosis of where model performance gaps arise across languages and writing systems.
Motivation
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.