• Home
  • Search
  • Explain and improve: LRP-inference fine-tuning for image captioning models
  • Cite Icon34
  • https://doi.org/10.1016/j.inffus.2021.07.008Copy DOI Icon

Explain and improve: LRP-inference fine-tuning for image captioning models

Show More
  • Abstract
  • Highlights & Summary
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This paper analyzes the predictions of image captioning models with attention mechanisms beyond visualizing the attention itself. We develop variants of Layer-wise Relevance Propagation (LRP) and gradient-based explanation methods, tailored to image captioning models with attention mechanisms. We compare the interpretability of attention heatmaps systematically against the explanations provided by explanation methods such as LRP, Grad-CAM, and Guided Grad-CAM. We show that explanation methods provide simultaneously pixel-wise image explanations (supporting and opposing pixels of the input image) and linguistic explanations (supporting and opposing words of the preceding sequence) for each word in the predicted captions. We demonstrate with extensive experiments that explanation methods (1) can reveal additional evidence used by the model to make decisions compared to attention; (2) correlate to object locations with high precision; (3) are helpful to “debug” the model, e.g. by analyzing the reasons for hallucinated object words. With the observed properties of explanations, we further design an LRP-inference fine-tuning strategy that reduces the issue of object hallucination in image captioning models, and meanwhile, maintains the sentence fluency. We conduct experiments with two widely used attention mechanisms: the adaptive attention mechanism calculated with the additive attention and the multi-head attention mechanism calculated with the scaled dot product.

Similar Papers
  • Research Article
  • Citations2

Explain ability and interpretability in machine learning models

  • Jan 01, 2020
  • Journal of Computer Science Applications and Information Technology
  • Vijay Kumar Adari +4
  • Research Article
  • Citations27

Interpretable Deep Learning for Pneumonia Detection Using Chest X-Ray Images

  • Jan 15, 2025
  • Information
  • Jovito Colin +1
  • Research Article

ClathPLM: Deep multi-view feature extraction with CNN and attention enhances clathrin protein identification.

  • May 01, 2026
  • Journal of molecular graphics & modelling
  • Shuxin Song +3
  • Research Article
  • Citations1

Layer-Wise Relevance Propagation Approach for Diagnosis of Drug-Naïve Men With Major Depressive Disorder Using Resting-State Electroencephalography

  • Jan 01, 2025
  • Depression and Anxiety
  • Eun-Gyoung Yi +3
  • Research Article
  • Citations3

Attention Heat Map-Based Black-Box Local Adversarial Attack for Synthetic Aperture Radar Target Recognition

  • Oct 01, 2024
  • Photogrammetric Engineering & Remote Sensing
  • Xuanshen Wan +3
  • Research Article

Quantifying Deep Learning Interpretability in Alzheimer's Disease Research: Validating Deep Learning Heatmaps through Meta‐Analysis

  • Dec 01, 2024
  • Alzheimer's & Dementia
  • Di Wang +5
  • Research Article
  • Citations47

Autism Spectrum Self-Stimulatory Behaviors Classification Using Explainable Temporal Coherency Deep Features and SVM Classifier

  • Jan 01, 2021
  • IEEE Access
  • Shuaibing Liang +3
  • Conference Article
  • Citations2

Image Denoising Algorithm Based on Multi-Scale Fusion and Adaptive Attention Mechanism

  • Feb 24, 2023
  • Zhitong Su +4
  • PDF
  • Research Article
  • Citations8

Object Detection Algorithm Based on Multiheaded Attention

  • May 02, 2019
  • Applied Sciences
  • Jie Jiang +3
  • Research Article
  • Citations1

Explainable AI for Cyber Threat Intelligence and Risk Assessment

  • Jan 01, 2020
  • Journal of Frontiers in Multidisciplinary Research
  • Ehimah Obuse +6
  • PDF
  • Research Article

Measured multi-source semi-supervised working condition recognition based on curvelet pooling and attention mechanism learning

  • Nov 14, 2025
  • Scientific Reports
  • Shuo Yang +3
  • Research Article
  • Citations8

Multi-Layer Perceptron Model Integrating Multi-Head Attention and Gating Mechanism for Global Navigation Satellite System Positioning Error Estimation

  • Jan 16, 2025
  • Remote Sensing
  • Xiuxun Liu +2
  • Conference Article

Operating Condition Diagnosis of Pumping Wells Based on 1DCNN-BiLSTM with Multi-Head Attention Mechanism

  • Aug 01, 2025
  • Tingfeng Hou +1
  • Research Article
  • Citations38

COVID-19 fake news detection: A hybrid CNN-BiLSTM-AM model

  • Jul 28, 2023
  • Technological Forecasting and Social Change
  • Huosong Xia +5
  • Research Article

Prediction of piRNA and Disease Association based on Graph Neural Network

  • Nov 27, 2025
  • International Journal of Biology and Life Sciences
  • Xiulian Fang
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.