Search papers, labs, and topics across Lattice.
This paper systematically evaluates the reliability of diffusion-based Large Vision-Language Models (dLVLMs) in comparison to autoregressive (AR) models, focusing on hallucination and bias across various dimensions. Key findings reveal that while dLVLMs reverse the yes-bias seen in AR models and maintain competitive hallucination rates, they suffer from significant accuracy collapse, particularly regarding underrepresented racial groups and in multiple-choice scenarios. The study highlights the unique mechanisms of diffusion generation that contribute to these reliability issues, emphasizing the need for careful consideration of generative paradigms and training data in model development.
dLVLMs not only reverse the yes-bias of AR models but also collapse in accuracy for underrepresented groups, revealing critical reliability gaps in diffusion-based approaches.
Diffusion-based Large Vision-Language Models (dLVLMs) have recently emerged as a compelling alternative to autoregressive (AR) LVLMs, offering advantages in parallel decoding, bidirectional context, and controllable generation. Despite rapid progress, their reliability properties remain largely uncharacterized. We present the first systematic reliability evaluation of hallucination and bias in dLVLMs, benchmarking six diffusion models against competitive AR baselines across four dimensions. Our key findings are: (1) dLVLMs reverse the yes-bias of AR models in binary visual queries; (2) they achieve competitive hallucination rates yet exhibit degraded linguistic quality; (3) they collapse to near-zero accuracy on underrepresented racial groups with opposite-polarity gender bias; and (4) they exhibit accuracy collapse in multiple-choice settings when the correct option is shorter than its distractors, associated with a length prior that emerges at the first denoising step. Tokens committed at late denoising steps with low confidence further correlate with hallucinated content, pointing to a mechanistic signal unique to diffusion generation. These patterns vary across model families, suggesting reliability is shaped by the generative paradigm together with training data.