Search papers, labs, and topics across Lattice.
The ImageEval 2026 shared task focuses on culturally grounded Arabic multimodal evaluation, featuring two main tasks: AynVQA, which assesses spoken visual question answering and image-grounded hallucination detection, and CRAI-Bench, which evaluates the cultural accuracy of text-to-image generation. With participation from 14 teams employing diverse methodologies such as zero-shot prompting and fine-tuning of vision-language models, the task reveals significant challenges in multimodal evaluation specific to Arabic contexts. The results underscore the need for culturally aware evaluation metrics and provide valuable datasets and evaluation scripts for future research in this area.
Culturally grounded multimodal evaluation reveals significant gaps in current AI systems' understanding of Arabic contexts, with implications for both accuracy and representation.
We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic (MSA), and (ii) CRAI-Bench, evaluating the cultural accuracy of text-to-image generation. A total of 14 teams participated in the test phase, with 12 teams submitting system description papers. Participating systems used a range of approaches, including zero-shot prompting, fine-tuning of vision-language models, speech-recognition pipelines, ensembling, and score calibration. We describe the task setup, datasets, evaluation procedure, and participating systems, and summarize the main results across the different tracks. All datasets and evaluation scripts from the shared task are released to the research community. The shared task highlights the challenges of culturally grounded multimodal evaluation, particularly for Arabic speech and image-text reasoning.