Benchmark Radar
AI BENCHMARK PROFILE

QuechuaTok

General AICoding & Software Engineering

A comparison of tokenization strategies (BPE, Unigram LM, WordPiece, PRPE) for Southern Quechua using a 200k-sentence corpus and a finite-state morphological analyzer as reference, with metrics including fertility rate, OOV rate, and morphological boundary accuracy.

Released
2026-06-22
Readiness
Paper only
Primary field
General AI

Why it matters

Evaluates tokenizer quality for agglutinative low-resource languages, showing that fertility rate alone is insufficient and that morphological boundary accuracy provides a more meaningful signal.

Motivation

Tokenization is a foundational step in NLP pipelines, yet standard evaluation metrics such as fertility rate fail to capture morphological correctness for agglutinative languages.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.