• Home
  • Search
  • Image caption generator using deep learning architecture with LSTM and their comparative analysis
  • https://doi.org/10.1049/icp.2025.4693Copy DOI Icon

Image caption generator using deep learning architecture with LSTM and their comparative analysis

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Access to precise visual information is vital for digital accessibility, particularly for blind and visuall y impaired individuals. Existing image captioning models still generate correct captions that make sense only when trained on tidy curated datasets, which is a constraint for real applications. The variability in this is problematic for assistive technolog ies and thus, consistent and reliable generation of captions is of critical importance. To overcome these issues, we used and compared three deep learning models VGG16, ResNet50 and EfficientNetB2 combined with an LSTM language model for automatic image captioning. From these, the EfficientNetB2 + LSTM outperformed the rest as it had the best optimized features and was the most efficient architecture that provided the best answer to caption reliability problem. They report BLEU-4 and ROUGE scores of 0.26 and 0.59, respectively, as well as a CIDEr score of 0.89 on the Flickr30k dataset. This demonstrates the objective performance of the model. On a subjective level, the captions produced were logical and contextually appropriate, and showed great promise as a means to improve digital access. Our findings indicate that EfficientNetB2 + LSTM can be successfully deployed in assistive technologies as it is the most efficient and effective model we have tested.

Similar Papers
  • Research Article

STEERING ANGLE DETECTION IN AUTONOMOUS CARS USING DEEP LEARNING

  • Jan 01, 2024
  • International Journal Of Trendy Research In Engineering And Technology
  • Aishwarya N +4
  • Research Article
  • Citations248

Deep learning based classification of breast tumors with shear-wave elastography

  • Aug 06, 2016
  • Ultrasonics
  • Qi Zhang +6
  • Research Article
  • Citations2

Smoke detection from foggy environment based on color spaces

  • Sep 30, 2021
  • International Journal of Applied Mathematics Electronics and Computers
  • Mehmet Erdal Özbek +1
  • Book Chapter
  • Citations10

Generation of Image Captions Using VGG and ResNet CNN Models Cascaded with RNN Approach

  • Jan 01, 2020
  • Madhuri Bhalekar +3
  • Research Article
  • Citations2

Towards high-performance deep learning architecture and hardware accelerator design for robust analysis in diffuse correlation spectroscopy

  • Oct 28, 2024
  • Computer Methods and Programs in Biomedicine
  • Zhenya Zang +6
  • Research Article

AI-Based Emotion Recognition in Education: Progress, Applications, and Open Challenges

  • Feb 05, 2026
  • Journal of Electrical Engineering and Computer Science (JEECS) | E-ISSN : 3089-5952
  • Rexcharles Enyinna Donatus +4
  • Conference Article

Arabic Legal Text Classification Using Pre-Trained Transformers and Deep Learning

  • Jun 19, 2025
  • Oussama Tahtah +2
  • Research Article
  • Citations23

Quantification of MR spectra by deep learning in an idealized setting: Investigation of forms of input, network architectures, optimization by ensembles of networks, and training bias.

  • Dec 19, 2022
  • Magnetic Resonance in Medicine
  • Rudy Rizzo +4
  • PDF
  • Research Article
  • Citations31

Comparative Study of Popular Deep Learning Models for Machining Roughness Classification Using Sound and Force Signals.

  • Nov 29, 2021
  • Micromachines
  • Binayak Bhandari
  • Conference Article
  • Citations1

MR-radiomic biopsy for estimation of malignancy grade in parotid gland cancer

  • Mar 02, 2020
  • Hidemi Kamezawa +3
  • Research Article
  • Citations16

A novel automatic image caption generation using bidirectional long-short term memory framework

  • Apr 19, 2021
  • Multimedia Tools and Applications
  • Zhongfu Ye +3
  • Book Chapter
  • Citations7

Image Caption Generation Using Neural Network Models and LSTM Hierarchical Structure

  • Sep 05, 2021
  • Prachi M Waghmare +1
  • PDF
  • Research Article
  • Citations122

An Overview of Image Caption Generation Methods

  • Jan 09, 2020
  • Computational Intelligence and Neuroscience
  • Haoran Wang +2
  • Research Article
  • Citations6

Enhancing Cross-Linguistic Image Caption Generation with Indian Multilingual Voice Interfaces using Deep Learning Techniques

  • Jan 01, 2024
  • Procedia Computer Science
  • Vijay A Sangolgi +5
  • PDF
  • Research Article
  • Citations7

Sequential Dual Attention: Coarse-to-Fine-Grained Hierarchical Generation for Image Captioning

  • Nov 12, 2018
  • Symmetry
  • Zhibin Guan +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.