AI BENCHMARK PROFILE
QuechuaTok
A comparison of tokenization strategies (BPE, Unigram LM, WordPiece, PRPE) for Southern Quechua using a 200k-sentence corpus and a finite-state morphological analyzer as reference, with metrics including fertility rate, OOV rate, and morphological boundary accuracy.
- Released
- 2026-06-22
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Evaluates tokenizer quality for agglutinative low-resource languages, showing that fertility rate alone is insufficient and that morphological boundary accuracy provides a more meaningful signal.
Motivation
Tokenization is a foundational step in NLP pipelines, yet standard evaluation metrics such as fertility rate fail to capture morphological correctness for agglutinative languages.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.