Benchmark Radar
AI BENCHMARK PROFILE

NOLLI

General AIKnowledge & ReasoningHAE-RAE

NOLLI is a procedurally generated English-Korean puzzle benchmark with 15 puzzle types (25 tasks, 7,500 items). Each instance is seed-regenerable, verified to have a unique solution, and scored deterministically. Difficulty is calibrated behaviorally to target accuracy bands.

Released
2026-08-05
Readiness
Runnable
Primary field
General AI

Why it matters

NOLLI addresses the lack of controlled cross-lingual benchmarks that separate presentation language from reasoning difficulty, enabling diagnosis of where model performance gaps arise across languages and writing systems.

Motivation

We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.