PACUTE
PACUTE is a diagnostic benchmark of 4,600 tasks evaluating morphological understanding in Filipino, covering six compositional levels from morpheme decomposition to syllabification.
- Released
- 2026-06-13
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Standard tokenizers obscure character-level and morphological structure, particularly for languages with non-concatenative morphology. This benchmark localizes where morphological understanding breaks down, distinguishing character access from productive composition, which informs tokenizer and model design for low-resource languages.
Motivation
Large language models (LLMs) process text as sequences of subword tokens, which can obscure the character-level and morphological structure that underlies word formation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.