Search papers, labs, and topics across Lattice.
This paper introduces SeVeR, a selective visual exposure framework designed to enhance 3D medical image question answering (VQA) by addressing the challenges posed by redundant visual token sequences in multi-sequence MRI data. The authors developed BreMRIs-VQA, a benchmark consisting of 1.19 million question-answer pairs, to evaluate the effectiveness of their approach, which employs modality-wise prototypes and change-aware gated attention to optimize visual retrieval. Experimental results demonstrate that SeVeR significantly improves both discriminative and generative performance while reducing the number of visual tokens processed during decoding.
Reducing visual token exposure by leveraging selective retrieval can dramatically enhance the performance of 3D medical image question answering systems.
Volumetric medical VQA requires reasoning over long and redundant 3D visual token sequences, especially in multi-sequence MRI where complementary modalities provide diverse diagnostic cues but expose the decoder to many repeated anatomical regions. To investigate reasoning under multi-sequence visual redundancy, we first introduce BreMRIs-VQA, a clinically curated breast MRI benchmark with 1.19M QA pairs from 71.0K sequences and 12.9K patients, covering both free-text and multiple-choice questions. We further propose SeVeR, a selective visual exposure framework that compresses dense volumes into modality-wise prototypes and retrieves complementary multi-level evidence with change-aware gated attention during decoding, trained with a marginal-utility self-consistency objective that suppresses unhelpful retrieval. Experiments on BreMRIs-VQA and public benchmarks show that SeVeR improves both discriminative and generative performance while exposing substantially fewer visual tokens.