• Home
  • Search
  • A multimedia content identification system using audio and visual integrated features
  • https://doi.org/10.1109/mwscas.2005.1594485Copy DOI Icon

A multimedia content identification system using audio and visual integrated features

  • Jan 1, 2005
  • Chih-Chang Chen +5 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

With an overflow of multimedia information around us and an urgent need to identify data accurately, an audio and visual identification system with a high accuracy rate is developed to meet the demand. Classification and feature extraction are performed separately on audio and visual signals. Pending on the temporal correlation of the feature vectors of objects and speakers, indexes of all objects included in an audio/visual sequence are listed in a time sequence. In integrating the audio/visual features, every object or character of the key frames has a set of feature vectors; the user can select and search specific characters that have the audio and visual features from the entire index set. Due to integrating the audio/visual identification results in the time order, the proposed identification system can increase the accuracy about 4% and 6% in our experiments, comparing with the results using the audio features and visual features separately, respectively.

Similar Papers
  • Conference Article
  • Citations9

Comparison of early and late fusion techniques for movie trailer genre labelling

  • Jul 01, 2020
  • J.H Mervitz +3
  • Research Article
  • Citations54

Recognition of isolated words using Zernike and MFCC features for audio visual speech recognition

  • Oct 21, 2014
  • International Journal of Speech Technology
  • Prashant Borde +3
  • Conference Article
  • Citations8

Discrimination comparison between audio and visual features

  • Nov 01, 2012
  • Chao Sui +3
  • Conference Article
  • Citations10

Audio-Visual Speech Recognition System Using Recurrent Neural Network

  • Oct 01, 2019
  • Yeh-Huann Goh +2
  • Conference Article
  • Citations10

Bimodal log-linear regression for fusion of audio and visual features

  • Oct 21, 2013
  • Ognjen Rudovic +2
  • Research Article
  • Citations42

Audio-Visual Event Localization by Learning Spatial and Semantic Co-Attention

  • Jan 01, 2023
  • IEEE Transactions on Multimedia
  • Cheng Xue +4
  • PDF
  • Research Article
  • Citations48

Audio-Visual Speech Recognition Using Lip Information Extracted from Side-Face Images

  • Jan 01, 2007
  • EURASIP Journal on Audio, Speech, and Music Processing
  • Koji Iwano +3
  • Book Chapter
  • Citations11

Multimodal PLSA for Movie Genre Classification

  • Jan 01, 2015
  • Hao-Zhi Hong +1
  • Book Chapter
  • Citations11

Chapter 21 - Optical flow-based representation for video action detection

  • Dec 12, 2014
  • Emerging Trends in Image Processing, Computer Vision, and Pattern Recognition
  • Samet Akpınar +1
  • Conference Article
  • Citations6

SAAMEAT

  • Oct 02, 2018
  • Fasih Haider +3
  • Research Article
  • Citations1

CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization

  • Mar 14, 2026
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Jinxing Zhou +5
  • Book Chapter
  • Citations2

Speech-Driven Facial Animation Using Manifold Relevance Determination

  • Jan 01, 2016
  • Samia Dawood +2
  • Conference Article
  • Citations41

Fast and robust search method for short video clips from large video collection

  • Aug 23, 2004
  • Junsong Yuan +2
  • Conference Article
  • Citations4

Predicting Conversation Outcomes Using Multimodal Transformer

  • Jul 18, 2021
  • Can Li +4
  • Research Article
  • Citations23

Investigation of acoustic and visual features for pig cough classification

  • Jun 01, 2022
  • Biosystems Engineering
  • Nan Ji +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.