• Home
  • Search
  • Efficient Channel Attention Based Encoder–Decoder Approach for Image Captioning in Hindi
  • Cite Icon18
  • https://doi.org/10.1145/3483597Copy DOI Icon

Efficient Channel Attention Based Encoder–Decoder Approach for Image Captioning in Hindi

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Image captioning refers to the process of generating a textual description that describes objects and activities present in a given image. It connects two fields of artificial intelligence, computer vision, and natural language processing. Computer vision and natural language processing deal with image understanding and language modeling, respectively. In the existing literature, most of the works have been carried out for image captioning in the English language. This article presents a novel method for image captioning in the Hindi language using encoder–decoder based deep learning architecture with efficient channel attention. The key contribution of this work is the deployment of an efficient channel attention mechanism with bahdanau attention and a gated recurrent unit for developing an image captioning model in the Hindi language. Color images usually consist of three channels, namely red, green, and blue. The channel attention mechanism focuses on an image’s important channel while performing the convolution, which is basically to assign higher importance to specific channels over others. The channel attention mechanism has been shown to have great potential for improving the efficiency of deep convolution neural networks (CNNs). The proposed encoder–decoder architecture utilizes the recently introduced ECA-NET CNN to integrate the channel attention mechanism. Hindi is the fourth most spoken language globally, widely spoken in India and South Asia; it is India’s official language. By translating the well-known MSCOCO dataset from English to Hindi, a dataset for image captioning in Hindi is manually created. The efficiency of the proposed method is compared with other baselines in terms of Bilingual Evaluation Understudy (BLEU) scores, and the results obtained illustrate that the method proposed outperforms other baselines. The proposed method has attained improvements of 0.59%, 2.51%, 4.38%, and 3.30% in terms of BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores, respectively, with respect to the state-of-the-art. Qualities of the generated captions are further assessed manually in terms of adequacy and fluency to illustrate the proposed method’s efficacy.

Similar Papers
  • Research Article
  • Citations31

A Deep Attention based Framework for Image Caption Generation in Hindi Language

  • Oct 07, 2019
  • Computación y Sistemas
  • Rijul Dhir +3
  • PDF
  • Research Article
  • Citations31

EFFNet-CA: An Efficient Driver Distraction Detection Based on Multiscale Features Extractions and Channel Attention Mechanism

  • Apr 08, 2023
  • Sensors (Basel, Switzerland)
  • Taimoor Khan +2
  • Research Article
  • Citations44

A Hindi Image Caption Generation Framework Using Deep Learning

  • Mar 15, 2021
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Santosh Kumar Mishra +3
  • Conference Article
  • Citations3

Novel Image Caption System Using Deep Convolutional Neural Networks (VGG16)

  • Jun 09, 2022
  • Alaa Sabeeh Salim +5
  • PDF
  • Research Article
  • Citations122

An Overview of Image Caption Generation Methods

  • Jan 09, 2020
  • Computational Intelligence and Neuroscience
  • Haoran Wang +2
  • Research Article
  • Citations15

Image Captioning using Artificial Intelligence

  • Apr 01, 2021
  • Journal of Physics: Conference Series
  • Yajush Pratap Singh +4
  • Research Article
  • Citations5

Generating image captions through multimodal embedding

  • May 14, 2019
  • Journal of Intelligent & Fuzzy Systems
  • Sandeep Kumar Dash +3
  • PDF
  • Research Article
  • Citations4

Emotion recognition model based on CLSTM and channel attention mechanism

  • Jan 01, 2022
  • ITM Web of Conferences
  • Yuxia Chen +2
  • PDF
  • Research Article
  • Citations5

High-level and Low-level Feature Set for Image Caption Generation with Optimized Convolutional Neural Network

  • Dec 29, 2022
  • Journal of Telecommunications and Information Technology
  • Roshni Padate +3
  • Research Article
  • Citations44

Boosting convolutional image captioning with semantic content and visual relationship

  • Aug 23, 2021
  • Displays
  • Cong Bai +4
  • Research Article
  • Citations60

Integration of textual cues for fine-grained image captioning using deep CNN and LSTM

  • Oct 19, 2019
  • Neural Computing and Applications
  • Neeraj Gupta +1
  • PDF
  • Research Article
  • Citations1

Oppositional Harris Hawks Optimization with Deep Learning-Based Image Captioning

  • Jan 01, 2023
  • Computer Systems Science and Engineering
  • V R Kavitha +6
  • Research Article
  • Citations11

A Benchmark for Feature-injection Architectures in Image Captioning

  • Dec 06, 2021
  • European Journal of Science and Technology
  • Rumeysa Keski̇n +4
  • Book Chapter

Advancing image captioning: Multilingual capabilities and Hindi-specific applications using transformer-based architectures with real-time smart focus camera for visually impaired users

  • Mar 05, 2026
  • Bijesh Thomas +3
  • Research Article
  • Citations12

Multi-view pedestrian captioning with an attention topic CNN model

  • Feb 08, 2018
  • Computers in Industry
  • Quan Liu +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.