Search papers, labs, and topics across Lattice.
This paper introduces EDRAC, the first comprehensive benchmark for machine reading comprehension (MRC) and generative question answering (QA) in dialectal Arabic, addressing the significant resource gap compared to Modern Standard Arabic. By leveraging a collaborative pipeline that combines human input and LLM evaluation, EDRAC encompasses 499 passages and nearly 5,000 QA pairs across five major Arabic dialects. The findings indicate critical discrepancies between the semantic quality of answers and their dialectal accuracy, underscoring the inadequacies of current evaluation metrics for dialectal Arabic NLP tasks.
EDRAC reveals that existing Arabic LLMs struggle with dialectal fidelity, exposing a critical gap in current benchmarks for Arabic NLP.
Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-choice QA, with limited coverage of naturally spoken dialects. Here, we aim to bridge this gap. We introduce EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic. EDRAC contains 499 passages derived from naturally occurring spoken interactions and 4,977 corresponding QA pairs generated through a human--LLM collaborative pipeline combining iterative generation, LLM-as-a-judge evaluation, and human verification. We benchmark Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics. Our results reveal substantial gaps between semantic answer quality and dialectal fidelity, highlighting the limitations of existing evaluation metrics for dialectal Arabic generation. EDRAC provides a realistic and challenging MRC benchmark for future research on dialectal Arabic NLP.