Search papers, labs, and topics across Lattice.
The SHROOM-Visions 2026 shared task focused on detecting hallucinations in large vision-language models, building on the SHEEP dataset to evaluate model performance across multiple languages. This iteration attracted significant participation, with 27 teams submitting over 600 systems, showcasing a competitive landscape in hallucination detection. Key results indicated that the top-performing systems achieved average scores of 0.58 in character-level correlation and 0.51 in intersection-over-union, significantly surpassing baseline performances by 30-40 points.
The top systems in hallucination detection outperformed baselines by up to 40 points, revealing significant advancements in tackling errors in vision-language models.
In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes in \textbf{Vision} language model\textbf{s}), which is hosted at the UncertaiNLP Workshop co-located with EMNLP 2026. Following the success of the 2024 and 2025 tasks, this time we aim to tackle hallucinations through a model-agnostic detection task focused on large vision-language models. Building on the recently introduced SHEEP dataset, designed for long-term evaluation across model generations, the task invites participants to detect and classify fine-grained hallucination spans in image-conditioned text generation (VQA, image captioning, etc.). The evaluation uses a five-class taxonomy of hallucinations spanning four languages: Chinese, English, French, and Italian. The shared task generated strong interest in the NLP community worldwide, with 27 teams contributing 600+ system submissions. The best systems achieve average scores of 0.58 in character-level correlation, 0.46 in label-conditioned correlation, and 0.51 in intersection-over-union (IoU) across four languages, outperforming the baselines by 30-40 points.