AI BENCHMARK PROFILE
VialectBench
Evaluates LLM robustness to Vietnamese dialectal rewrites across emotion recognition, natural language inference, question answering, and multiple-choice QA, measuring performance degradation across six dialect groups.
- Released
- 2026-08-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Highlights that model performance on standard Vietnamese does not guarantee reliable behavior under regional variation, informing deployment decisions for Vietnamese-language applications.
Motivation
Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.