• Home
  • Search
  • Text-Conditional Visual-Language Alignment for Video Captioning
  • https://doi.org/10.1109/tcsvt.2025.3616201Copy DOI Icon

Text-Conditional Visual-Language Alignment for Video Captioning

  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Video captioning remains a challenging task due to the diverse video content and the complex relationships between visual and textual elements. Recent efforts predominantly focus on multimodal architecture designs trained with paired video-caption data. Nonetheless, the learning paradigm suffers from the “one-to-many” corresponding problem, since one source video is mapped to multiple caption annotations. The difficulty of video captioning is further exacerbated by the poor-written captions, which mislead the captioner with irrelevant information. Essentially, the problem stems from the inadequate alignment between video and caption. In this work, we propose a Text-Conditional Alignment Transformer, which fully exploits the rich information provided by diverse labeled captions, and avoids the impacts of label ambiguity and noise. To alleviate the challenge of the “one-to-many” correspondence, we introduce Text-conditioned Video Encoding, which diversifies the video representation by emphasizing the spatial-temporal visual areas relevant to the given descriptions while filtering out redundant visual information. The refined video representation is well-aligned to match the corresponding text description, and naturally converts the “one-to-many” mapping to “one-to-one” mapping. To deal with the noisy annotations, we propose Quality-aware Caption Decoding. We first dynamically measure the qualities of different captions corresponding to the same video in a reference-free manner. Then the estimated qualities are further utilized as auxiliary signals, guiding the model to perform quality-aligned learning from noisy captions. We conduct extensive experiments on MSR-VTT, MSVD, VATEX and ActivityNet-Entities datasets, and demonstrate their consistent performance improvements compared to state-of-the-arts.

Similar Papers
  • Research Article
  • Citations4

Implicit and explicit commonsense for multi-sentence video captioning

  • Jul 05, 2024
  • Computer Vision and Image Understanding
  • Shih-Han Chou +2
  • Research Article
  • Citations15

Adaptive Curriculum Learning for Video Captioning

  • Jan 01, 2022
  • IEEE Access
  • Shanhao Li +2
  • Conference Article

The Video Captioning Method Based On The Spatial- Temporal Information and Attention Mechanism

  • Aug 17, 2021
  • Ou Ye +5
  • Research Article
  • Citations108

Video Captioning Using Global-Local Representation.

  • Oct 01, 2022
  • IEEE Transactions on Circuits and Systems for Video Technology
  • Liqi Yan +6
  • Research Article

Arthritis websites — Deconstructing the messages to identify strong versus weak performers

  • Dec 01, 2002
  • Journal of Medical Marketing
  • H R Moskowitz +4
  • Research Article
  • Citations1

The poetics of the rebus: Word, image and the dynamics of reading in the poster of the 1920s and 1930s

  • Jul 01, 1997
  • Word & Image
  • David Scott
  • Conference Article
  • Citations5

Beyond verbs: Understanding actions in videos with text

  • Dec 01, 2016
  • Shujon Naha +1
  • Conference Article
  • Citations7

Boosting Video Representation Learning with Multi-Faceted Integration

  • Jun 01, 2021
  • Zhaofan Qiu +5
  • Conference Article
  • Citations565

Learning from Noisy Large-Scale Datasets with Minimal Supervision

  • Jul 01, 2017
  • Andreas Veit +5
  • Research Article

Empowering Educational Researchers through Effective Data Presentation Methods: Balancing Clarity, Accuracy, and Interpretability

  • Jan 19, 2026
  • ICoBITS
  • Rahmahdalena Putri Khairunnisa +1
  • Research Article

Adaptive Dual Video Summarization: From Dynamic Keyframes to Captions

  • Jan 01, 2025
  • IEEE Transactions on Multimedia
  • Zhenzhen Hu +6
  • Research Article

Learning Procedural-Aware Video Representations Through State-Grounded Hierarchy Unfolding

  • Mar 14, 2026
  • Jinghan Zhao +2
  • PDF
  • Research Article
  • Citations9

Deep Learning-Based Context-Aware Video Content Analysis on IoT Devices

  • Jun 04, 2022
  • Electronics
  • Gad Gad +4
  • Conference Article
  • Citations4

Natural Language Description for Videos Using NetVLAD and Attentional LSTM

  • Jun 01, 2020
  • V K Jeevitha +1
  • Research Article
  • Citations5

Content of YouTube videos on cassava production and processing in Nigeria

  • Nov 11, 2021
  • Journal of Agricultural Extension
  • Tajudeen Oyekunle Amoo Banmeke +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.