Search papers, labs, and topics across Lattice.
This study evaluates the adversarial robustness of five leading Arabic Language Models against various adversarial attack strategies, revealing significant vulnerabilities in their performance. Notably, the insertion of diacritics can lead to a staggering 92% accuracy drop in some models, while word-level manipulations and sentence-level paraphrasing also yield substantial performance degradation. Although adversarial training enhances model resilience, challenges remain, particularly against character-level noise, underscoring the complexities of securing morphologically rich languages like Arabic.
Arabic Language Models can suffer up to a 92% accuracy drop from simple diacritic insertions, revealing critical vulnerabilities in their robustness.
The emergence of the recent outstanding capabilities of Arabic Language Models has opened doors for exposing their vulnerabilities. One of the major security risks associated with such Natural Language Processing models is adversarial attacks. These attacks can deceive the model into the wrong prediction, raising critical model security and safety concerns. This study aims to assess the robustness of five state-of-the-art Arabic Language Models under a distinct set of Arabic adversarial attacks applied at various levels of granularity and using different example generation strategies. We also explore a defense technique based on adversarial training to enhance model robustness. The results show that insertion of diacritics can reduce the accuracy of some models by 92% while maintaining a low perturbation distance. For word-level attacks, manipulating Arabic conjunctions preserves high semantic similarity scores, low perturbation distance, and leads to an accuracy degradation of up to 58%. For sentence-level attacks, paraphrasing proves its effectiveness by an average reduction of 76% in the victim models'performance. While adversarial training improves overall resilience, with MARBERT being the most robust and AraBERT showing the greatest relative gains, challenges persist, particularly against character-level noise. These findings highlight both the potential and limitations of current defense strategies in morphologically rich languages like Arabic.