Search papers, labs, and topics across Lattice.
This paper introduces EEG-EditBench, a diagnostic benchmark designed to probe the visual information utilized by EEG-to-image retrieval models through controlled image edits. By evaluating eight EEG visual decoding models on 2,137 quality-controlled edits derived from the 200 THINGS-EEG2 dataset, the study uncovers that high aggregate retrieval accuracy does not necessarily imply robustness to fine-grained changes in visual attributes. The findings highlight significant discrepancies between standard retrieval performance and the models' ability to maintain sensitivity to specific visual modifications, revealing critical insights into model behavior that are often obscured by overall accuracy metrics.
Strong EEG-to-image retrieval models falter when faced with subtle visual edits, challenging assumptions about their robustness.
Recent EEG-to-image retrieval models have achieved strong performance in identifying viewed images from semantically diverse candidates. Yet such success does not reveal what visual information supports the match. A model may readily identify a cheetah among tools, plants, and vehicles, but can it still distinguish the viewed cheetah from the same scene with the cheetah replaced by a dog? Motivated by this question, we introduce EEG-EditBench, a diagnostic benchmark that examines this question through controlled edits of object identity, attributes, background, and object presence. Built from the 200 THINGS-EEG2 test images, EEG-EditBench contains 2,137 quality-controlled edits and evaluates eight representative EEG visual decoding models. Our results show that strong standard retrieval does not consistently transfer to edit-based evaluation, with fine-grained attribute changes presenting the greatest challenge. EEG-EditBench reveals model behavior hidden by aggregate retrieval accuracy and provides a controlled basis for studying what visual information EEG-image models preserve. The code and complete dataset are publicly available.