Search papers, labs, and topics across Lattice.
This paper introduces Perception-Enhanced Alignment Direct Preference Optimization (PEA-DPO), a novel framework designed to improve the alignment of multimodal large language models (LLMs) with human preferences by addressing the issue of visual insensitivity. Through a thorough representational analysis, the authors identify two specific failure modes鈥擜cross-Image Insensitivity and Within-Image Insensitivity鈥攖hat hinder effective multimodal preference optimization. Empirical evaluations reveal that PEA-DPO not only enhances sensitivity to visual context but also significantly reduces hallucinations while maintaining the language modeling capabilities of the base model across various scales.
Visual insensitivity in multimodal LLMs can be effectively mitigated with a new framework that enhances alignment and reduces hallucinations.
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal settings remains unexplored. Through representational analysis, we identify a key limitation in multimodal preference optimization, which we term visual insensitivity: models often fail to distinguish between images and those with critical visual context removed. Our theoretical analysis further uncovers two manifestations of this problem, namely Across-Image Insensitivity and Within-Image Insensitivity. To address these challenges, we propose Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity. We further provide a theoretical analysis demonstrating that PEA-DPO provably mitigates both failure modes. Empirical results demonstrate that PEA-DPO enhances sensitivity to visual context while preserving the language modeling capacity of the base model. Evaluations across three hallucination benchmarks using MLLMs of varying scales show that PEA-DPO effectively mitigates visual insensitivity, achieves stronger multimodal alignment, and substantially reduces hallucinations.