• Home
  • Search
  • Improving Narrative Coherence in Dense Video Captioning through Transformer and Large Language Models
  • Cite Icon2
  • https://doi.org/10.36548/jiip.2025.2.005Copy DOI Icon

Improving Narrative Coherence in Dense Video Captioning through Transformer and Large Language Models

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Dense video captioning aims to identify events within a video and generate natural language descriptions for each event. Most existing approaches adhere to a two-stage framework consisting of an event proposal module and a caption generation module. Previous methodologies have predominantly employed convolutional neural networks and sequential models to describe individual events in isolation. However, these methods limit the influence of neighboring events when generating captions for a specific segment, often resulting in descriptions that lack coherence with the broader storyline of the video. To address this limitation, we propose a captioning module that leverages both Transformer architecture and a Large Language Model (LLM). A convolutional and LSTM-based proposal module is used to detect and localize events within the video. An encoder-decoder-based Transformer model generates an initial caption for each proposed event. Additionally, we introduce a Large Language Model (LLM) that takes the set of individually generated event captions as input and produces a coherent, multi-sentence summary. This summary captures cross-event dependencies and provides a contextually unified and narratively rich description of the entire video. Extensive experiments on the ActivityNet dataset demonstrate that the proposed model, Transformer-LLM based Dense Video Captioning (TL-DVC), achieves a 9.22% improvement over state-of-the-art models, increasing the Meteor score from 11.28 to 12.32.

Similar Papers
  • Research Article

Evaluating gpt-4 for zero-shot classification of bleeding and clotting events: Can large language models serve as second reviewers?

  • Nov 03, 2025
  • Blood
  • Samantha Rizzo +5
  • Research Article
  • Citations1

Guardians of digital safety: benchmarking large language models in the fight against online toxicity

  • Dec 13, 2025
  • Journal of Big Data
  • Nouar Aldahoul +3
  • Research Article
  • Citations65

Model tuning or prompt Tuning? a study of large language models for clinical concept and relation extraction

  • Mar 26, 2024
  • Journal of Biomedical Informatics
  • Cheng Peng +6
  • Research Article
  • Citations2

Multi-objective evolutionary neural architecture search for medical image analysis using transformer and large language models in advancing public health

  • Jul 01, 2025
  • Applied Soft Computing
  • Yu Sun +3
  • Research Article

Quantitative Models with Qualitative Scenarios: Simulation with Transformer and Large Language Models

  • Jun 19, 2025
  • The Journal of Financial Data Science
  • Irene Aldridge +1
  • Conference Article
  • Citations2

Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models

  • Apr 05, 2025
  • Ahmad Antari +3
  • Research Article
  • Citations7

COFFE: A Code Efficiency Benchmark for Code Generation

  • Jun 19, 2025
  • Proceedings of the ACM on Software Engineering
  • Yun Peng +3
  • Conference Article

A Pre-training Method Inspired by Large Language Model for Power Named Entity Recognition

  • Nov 24, 2023
  • Qionglan Na +5
  • Preprint Article

From Tokens To Agents: A Researcher's Guide To Understanding Large Language Models

  • Feb 27, 2026
  • arXiv (Cornell University)
  • Daniele Barolo
  • Research Article

A systematic review on the generative AI applications in human medical genetics

  • Jan 20, 2026
  • Frontiers in Genetics
  • Anton Changalidis +3
  • Book Chapter

Explainable AI (XAI) for Classical ML and LLMs

  • May 16, 2026
  • Brijendra Parasnath Gupta +6
  • Conference Article
  • Citations1

Theoretical Insights into Fine-Tuning Attention Mechanism: Generalization and Optimization

  • Sep 01, 2025
  • Xinhao Yao +7
  • Research Article

LLM4Netlist: LLM-enabled Step-based Netlist Generation from Natural Language Description

  • Jan 01, 2025
  • IEEE Journal on Emerging and Selected Topics in Circuits and Systems
  • Kailiang Ye +6
  • Research Article
  • Citations2

Narrative Feature or Structured Feature? A Study of Large Language Models to Identify Cancer Patients at Risk of Heart Failure.

  • Jan 01, 2024
  • AMIA ... Annual Symposium proceedings. AMIA Symposium
  • Ziyi Chen +6
  • Research Article

R2GenCSR: Mining Contextual and Residual Information for LLMs-based Radiology Report Generation.

  • Jan 01, 2026
  • IEEE journal of biomedical and health informatics
  • Xiao Wang +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.