Search papers, labs, and topics across Lattice.
This paper introduces GUIDE, a novel framework designed to enhance multimodal models' adherence to language instructions regarding evidence usage during reasoning and generation. By integrating grouped parameter-efficient adaptation with instruction-conditioned gating, GUIDE effectively modulates how models utilize different evidence pathways, leading to a more structured and aligned evidence reliance. Experimental results across various datasets demonstrate that GUIDE not only improves robustness against targeted evidence perturbations but also allows for controllable modulation of evidence contributions, thereby extending the capabilities of multimodal instruction following.
GUIDE reshapes how multimodal models utilize evidence, enhancing their robustness and alignment with language instructions during reasoning tasks.
Large multimodal models follow instructions about what to generate, but not necessarily about what evidence to rely on. Hence, models may continue to depend on shortcut-associated cues even when instructions suggest otherwise. We introduce GUIDE, a framework for controlling internal evidence usage through language instructions. GUIDE combines grouped parameter-efficient adaptation with instruction-conditioned gating to modulate multimodal evidence pathways during reasoning and generation. We further introduce a pathway-level evaluation framework that characterizes instruction-conditioned evidence modulation through reliance sensitivity, controlled perturbation analysis, pathway modulation, and autoregressive decoding dynamics. Across multimodal reasoning, classification, and generation, GUIDE induces structured and instruction-aligned redistribution of evidence reliance while largely preserving task behavior. Experiments on GQA, TextVQA, MM-IMDb, CREMA-D, RAVDESS, and Flickr30K show that GUIDE improves robustness under targeted evidence perturbations and enables controllable modulation across diverse multimodal settings. This suggests that multimodal instruction following can extend beyond output control toward regulating how different evidence sources contribute to model predictions.