• Home
  • Search
  • CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
  • Cite Icon1
  • https://doi.org/10.1609/aaai.v40i16.38374Copy DOI Icon

CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization

Show More
  • Abstract
  • Literature Map
  • Citations
  • Similar Papers
Abstract

The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging weakly-supervised setting (W-DAVEL task), where only video-level event labels are provided and the temporal boundaries of each event are unknown. We address W-DAVEL by exploiting cross-modal salient anchors, which are defined as reliable timestamps that are well predicted under weak supervision and exhibit highly consistent event semantics across audio and visual modalities. Specifically, we propose a Mutual Event Agreement Evaluation module, which generates an agreement score by measuring the discrepancy between the predicted audio and visual event classes. Then, the agreement score is utilized in a Cross-modal Salient Anchor Identification module, which identifies the audio and visual anchor features through global-video and local temporal window identification mechanisms. The anchor features after multimodal integration are fed into an Anchor-based Temporal Propagation module to enhance event semantic encoding in the original temporal audio and visual features, facilitating better temporal localization under weak supervision. We establish benchmarks for W-DAVEL on both the UnAV-100 and ActivityNet1.3 datasets. Extensive experiments demonstrate that our method achieves state-of-the-art performance.

Similar Papers
  • Conference Article

A multimedia content identification system using audio and visual integrated features

  • Jan 01, 2005
  • Chih-Chang Chen +5
  • Research Article
  • Citations54

Recognition of isolated words using Zernike and MFCC features for audio visual speech recognition

  • Oct 21, 2014
  • International Journal of Speech Technology
  • Prashant Borde +3
  • Book Chapter
  • Citations11

Multimodal PLSA for Movie Genre Classification

  • Jan 01, 2015
  • Hao-Zhi Hong +1
  • Research Article
  • Citations3

Vision-language constraint graph representation learning for unsupervised vehicle re-identification

  • Jun 12, 2024
  • Expert Systems With Applications
  • Dong Wang +6
  • Conference Article
  • Citations21

METAL: Minimum Effort Temporal Activity Localization in Untrimmed Videos

  • Jun 01, 2020
  • Da Zhang +2
  • Conference Article
  • Citations8

Discrimination comparison between audio and visual features

  • Nov 01, 2012
  • Chao Sui +3
  • Conference Article
  • Citations9

Comparison of early and late fusion techniques for movie trailer genre labelling

  • Jul 01, 2020
  • J.H Mervitz +3
  • Research Article
  • Citations88

Multimodal sentiment analysis with unidirectional modality translation

  • Sep 20, 2021
  • Neurocomputing
  • Bo Yang +3
  • PDF
  • Research Article
  • Citations48

Audio-Visual Speech Recognition Using Lip Information Extracted from Side-Face Images

  • Jan 01, 2007
  • EURASIP Journal on Audio, Speech, and Music Processing
  • Koji Iwano +3
  • Conference Article
  • Citations10

Audio-Visual Speech Recognition System Using Recurrent Neural Network

  • Oct 01, 2019
  • Yeh-Huann Goh +2
  • Conference Article
  • Citations8

Learning a mid-level feature space for cross-media regularization

  • Jul 01, 2014
  • Yunchao Wei +4
  • Research Article
  • Citations42

Audio-Visual Event Localization by Learning Spatial and Semantic Co-Attention

  • Jan 01, 2023
  • IEEE Transactions on Multimedia
  • Cheng Xue +4
  • Book Chapter
  • Citations11

Chapter 21 - Optical flow-based representation for video action detection

  • Dec 12, 2014
  • Emerging Trends in Image Processing, Computer Vision, and Pattern Recognition
  • Samet Akpınar +1
  • Book Chapter
  • Citations2

Speech-Driven Facial Animation Using Manifold Relevance Determination

  • Jan 01, 2016
  • Samia Dawood +2
  • Conference Article
  • Citations41

Fast and robust search method for short video clips from large video collection

  • Aug 23, 2004
  • Junsong Yuan +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.