Search papers, labs, and topics across Lattice.
This paper introduces ArabicDialectSafety, a comprehensive dataset of 25,071 prompts annotated for safety classification across six Arabic dialects. The authors establish a dual-task evaluation framework that assesses both binary safety detection and detailed harm classification, revealing that fine-tuned MARBERTv2 significantly outperforms other models, including specialized LLMs, with Macro-F1 scores of 0.95 and 0.90, respectively. The findings highlight the importance of dialect conditioning at the representation level, while also exposing performance gaps in low-resource Maghrebi dialects, emphasizing the need for targeted improvements in these areas.
Fine-tuned MARBERTv2 outperforms leading LLMs in Arabic safety classification, revealing critical dialect-specific performance gaps.
We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestinian, and Moroccan. The dataset is annotated with dialect labels and seven fine-grained harm categories. We introduce a dual-task evaluation framework for binary safe/unsafe detection and granular harm classification across dialects. Benchmarking seven supervised and generative models, we find that fine-tuned MARBERTv2 achieves the strongest performance, with Macro-F1 scores of 0.95 for binary classification and 0.90 for granular classification, substantially outperforming prompted frontier LLMs, including Arabic-specialized models. Our analyses show that dialect conditioning is most effective when integrated at the representation level, while significant performance gaps remain for low-resource Maghrebi dialects. We further evaluate seven frontier LLMs as response generators on harmful dialectal Arabic prompts and observe unsafe generation rates below 5 percent across models. We release the dataset and code upon acceptance to support future research on dialect-aware Arabic safety evaluation. Warning: This paper contains examples of harmful and potentially offensive content included solely for research purposes.