Search papers, labs, and topics across Lattice.
To address severe performance degradations of multilingual LLMs on regional vernaculars, the authors construct 5-Dialects-BN, a multi-annotation benchmark spanning five regional varieties of Bangla: Chittagong, Barisal, Noakhali, Sylhet, and Rangpur. Standard multilingual evaluation predominantly overlooks non-standard dialects and Romanized transliterations despite Bangla being the world's sixth most spoken language. The resulting dataset provides 6,000 curated instances aligned across native dialect text, Romanized transliteration, Standard Bangla, English, and subjectivity labels to systematically benchmark tasks from normalization to LoRA fine-tuning.
Romanized and regional dialects represent a massive evaluation blind spot for frontier LLMs鈥攅ven across globally dominant low-resource languages like Bangla.
Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies this gap: existing resources overwhelmingly target Standard Bangla, leaving its regional dialects without the benchmarks needed to develop or evaluate dialect-aware systems. We address this gap with 5-Dialects-BN, the first multi-annotation Bangla dialect benchmark to align Romanized transliteration with dialectal text, Standard Bangla, English, and subjectivity labels across five regional varieties. The dataset comprises 6,000 manually annotated entries spanning five major dialects: Chittagong, Barisal, Noakhali, Sylhet, and Rangpur (Chittagong 1,900; Noakhali 1,500; Sylhet 1,200; Barisal 700; Rangpur 700), reflecting natural online availability. Each entry is enriched with five aligned annotations: the original dialectal text, a Romanized transliteration, an English translation, a Standard Bangla translation, and a subjectivity label (subjective vs. objective). Annotations were produced and cross-validated by native speakers and undergraduate linguistics students to ensure dialectal authenticity and semantic fidelity. The resulting resource supports a diverse suite of tasks, including dialect identification, dialect-to-standard normalization, machine translation, subjectivity classification, and parameter-efficient fine-tuning (e.g., LoRA) of multilingual LLMs. By providing a standardized, multi-annotation benchmark, 5-Dialects-BN enables principled evaluation of LLMs on dialectally diverse Bangla and lays a foundation for further research in low-resource, dialect-aware NLP.