- https://doi.org/10.1109/tmm.2026.3651099
Mitigating Hallucinations in Large Vision-Language Models via Visual-Enhanced Contrastive Decoding
- Jan 1, 2026
- IEEE Transactions on Multimedia
- Pengpeng Qiang +5 more
Despite significant advancements in large visual-language models (LVLMs), hallucinations remain a major bottleneck in their practical applications. One key factor contributing to hallucinations is the over-reliance on language priors during the autoregressive text generation process. Visual Contrastive Decoding (VCD), a popular technique for mitigating hallucinations, perturbs the visual input and compares the perturbed output with the original. However, it often overlooks the gradual attenuation of visual information within the decoder, limiting the model's ability to generate text based on actual visual content. We propose a novel, training-free method—Visual-Enhanced Contrastive Decoding (VECD)—which addresses this issue by amplifying visual information within the decoder, thereby reducing hallucinations caused by excessive reliance on language priors. VECD dynamically selects later layers for visual injection, while retaining only essential visual tokens in early layers. This approach enhances the generation process by adaptively balancing visual and language priors. By comparing outputs with and without visual amplification, we derive a refined probability distribution for the next token. Moreover, we improve the beam search algorithm by introducing a visually guided token selection strategy, enabling the generation of text that aligns more closely with the image content. Our extensive experiments show that VECD significantly reduces hallucinations and improves the quality of generated text, demonstrating its effectiveness as a practical solution.