• Home
  • Search
  • Progressive Video Summarization via Multimodal Self-supervised Learning
  • Open Access IconOpen Access
  • Cite Icon51
  • https://doi.org/10.1109/wacv56688.2023.00554Copy DOI Icon

Progressive Video Summarization via Multimodal Self-supervised Learning

  • Jan 1, 2023
  • Haopeng Li +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Modern video summarization methods are based on deep neural networks that require a large amount of annotated data for training. However, existing datasets for video summarization are small-scale, easily leading to over-fitting of the deep models. Considering that the annotation of large-scale datasets is time-consuming, we propose a multimodal self-supervised learning framework to obtain semantic representations of videos, which benefits the video summarization task. Specifically, the self-supervised learning is conducted by exploring the semantic consistency between the videos and text in both coarse-grained and fine-grained fashions, as well as recovering masked frames in the videos. The multimodal framework is trained on a newly-collected dataset that consists of video-text pairs. Additionally, we introduce a progressive video summarization method, where the important content in a video is pinpointed progressively to generate better summaries. Extensive experiments have proved the effectiveness and superiority of our method in rank correlation coefficients and F-score <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> .

Similar Papers
  • Conference Article
  • Citations48

Joint Video Summarization and Moment Localization by Cross-Task Sample Transfer

  • Jun 01, 2022
  • Hao Jiang +1
  • Research Article
  • Citations18

Efficient and Privacy Preserving Video Transmission in 5G-Enabled IoT Surveillance Networks: Current Challenges and Future Directions

  • Jan 05, 2021
  • IEEE Network
  • Khan Muhammad +5
  • PDF
  • Research Article
  • Citations67

To Compress or Not to Compress-Self-Supervised Learning and Information Theory: A Review.

  • Mar 12, 2024
  • Entropy
  • Ravid Shwartz Ziv +1
  • Research Article
  • Citations61

Dynamic graph convolutional network for multi-video summarization

  • Jun 20, 2020
  • Pattern Recognition
  • Jiaxin Wu +2
  • Research Article
  • Citations28

Improving quantitative MRI using self-supervised deep learning with model reinforcement: Demonstration for rapid T1 mapping.

  • Feb 11, 2024
  • Magnetic resonance in medicine
  • Wanyu Bian +2
  • Research Article
  • Citations45

Visual saliency models for summarization of diagnostic hysteroscopy videos in healthcare systems

  • Sep 06, 2016
  • SpringerPlus
  • Khan Muhammad +3
  • PDF
  • Research Article
  • Citations72

Self-Supervised Learning for Scene Classification in Remote Sensing: Current State of the Art and Perspectives

  • Aug 17, 2022
  • Remote Sensing
  • Paul Berg +2
  • Conference Article
  • Citations23

Self-Attention Recurrent Summarization Network with Reinforcement Learning for Video Summarization Task

  • Jul 05, 2021
  • Aniwat Phaphuangwittayakul +4
  • Research Article

Self-Supervised Deep Learning Models for Low-Resource NLP Applications

  • Feb 26, 2026
  • International Journal of Research and Review in Applied Science, Humanities, and Technology
  • Pardeep Kaur
  • Book Chapter
  • Citations7

Query Focused Video Summarization: A Review

  • Jan 01, 2022
  • Rakhi Akhare +1
  • Book Chapter
  • Citations1

Deep Neural Networks and Auditory Imagery

  • Nov 10, 2022
  • André Ofner +1
  • Book Chapter
  • Citations1

Extractive Text-Based Summarization of Arabic Videos: Issues, Approaches and Evaluations

  • Jan 01, 2019
  • Mohamed Amine Menacer +9
  • Conference Article
  • Citations8

A General Framework for Inverse Problem Solving using Self-Supervised Deep Learning: Validations in Ultrasound and Photoacoustic Image Reconstruction

  • Sep 11, 2021
  • Jingke Zhang +4
  • Conference Article
  • Citations46

Viewpoint-Aware Video Summarization

  • Jun 01, 2018
  • Atsushi Kanehira +3
  • Research Article

GARNN-AE-LSTM: A Multimodal Deep Learning Approach for High-Accuracy Video Summarization.

  • Oct 10, 2025
  • Journal of visualized experiments : JoVE
  • Jiasheng Jin +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.