Search papers, labs, and topics across Lattice.
This study systematically evaluates the robustness of Large Language Models (LLMs) to variations in Vietnamese dialects using a newly introduced benchmark, VialectBench, which includes 400 Standard Vietnamese instances and 2,400 dialectal rewrites across multiple tasks. The results reveal that dialectal inputs lead to an average performance drop of 2.82% across ten instruction-tuned models, with significant variability in robustness among different dialect groups. Notably, the Central dialect group exhibits the highest harmful-flip rate, indicating that proficiency in Standard Vietnamese does not ensure reliable performance in dialectal contexts.
Dialectal variations can degrade LLM performance by over 6% in some cases, revealing a critical gap in their robustness to real-world language use.
Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through dialect-to-standard normalization instead of measuring how the model fails under Vietnamese dialectal inputs. To address this gap, we present the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks, quantifying performance degradation and failure patterns. We introduce VialectBench (Vietnamese Dialects Benchmarking), a controlled benchmark for testing whether model decisions remain stable across six Vietnamese dialect groups. VialectBench contains 400 Standard Vietnamese source instances and 2,400 human-written dialectal rewrites spanning emotion recognition (ER), natural language inference (NLI), question answering (QA), and multiple-choice question answering (MCQA). Dataset evaluation with a fixed reference language model shows that the dialectal rewrites induce a measurable model-relative likelihood shift while remaining nearly equal in length to their Standard counterparts. Across ten instruction-tuned models, dialectal inputs reduce average performance by 2.82%, and no evaluated model is fully dialect-invariant. All four tasks are affected, with QA showing the largest average degradation. Robustness also varies substantially across dialect groups: PNT3 and PNT2 cause the largest average performance drops, at 6.17% and 4.73%, respectively, whereas PNB slightly improves average performance by 0.42%. The Central dialect group (PNT1-PNT4) also yields the highest average harmful-flip rate across all models, at 6.54%. These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation.