• Cite Icon2
  • https://doi.org/10.1145/3341162.3344861Copy DOI Icon

Audio-visual TED corpus

  • Sep 9, 2019
  • Guan-Lin Chao +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We present a variety of new visual features in extension to the TED-LIUM corpus. We re-aligned the original TED talk audio transcriptions with official TED.com videos. By utilizing state-of-the-art models for face and facial landmarks detection, optical character recognition, object detection and classification, we extract four new visual features that can be used for Large-Vocabulary Continuous Speech Recognition (LVCSR) systems, including facial images, landmarks, text, and objects in the scenes. The facial images and landmarks can be used in combination with audio for audio-visual acoustic modeling where the visual modality provides robust features in adverse acoustic environments. The contextual information, i.e. extracted text and detected objects in the scene can be used as prior knowledge to create contextual language models. Experimental results showed the efficacy of using visual features on top of acoustic features for speech recognition in overlapping speech scenarios.

Similar Papers
  • Conference Article
  • Citations13

Recent improvements of the SpeeD Romanian LVCSR system

  • May 01, 2014
  • Horia Cucu +4
  • PDF
  • Research Article
  • Citations7

Detecting Facial Region and Landmarks at Once via Deep Network †

  • Aug 09, 2021
  • Sensors (Basel, Switzerland)
  • Taehyung Kim +2
  • Research Article
  • Citations48

Acoustic models of the elderly for large‐vocabulary continuous speech recognition

  • Jun 09, 2004
  • Electronics and Communications in Japan (Part II: Electronics)
  • Akira Baba +4
  • Conference Article
  • Citations8

Language model cross adaptation for LVCSR system combination

  • Sep 26, 2010
  • Xunying Liu +2
  • Research Article
  • Citations22

Modelling Semantic Context of OOV Words in Large Vocabulary Continuous Speech Recognition

  • Feb 08, 2017
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Imran Sheikh +3
  • Conference Article

Large Vocabulary Continuous Audio-Visual Speech Recognition

  • Oct 02, 2018
  • George Sterpu
  • Research Article
  • Citations4

Construction and evaluation of language models based on stochastic context‐free grammar for speech recognition

  • Oct 23, 2002
  • Systems and Computers in Japan
  • Chiori Hori +3
  • Conference Article
  • Citations2

Large vocabulary continuous speech recognition based on cross-morpheme phonetic information

  • Oct 04, 2004
  • In-Jeong Choi +2
  • Research Article
  • Citations65

Combining Data-Driven and Model-Driven Methods for Robust Facial Landmark Detection

  • Feb 12, 2018
  • IEEE Transactions on Information Forensics and Security
  • Hongwen Zhang +3
  • Conference Article
  • Citations88

Constrained Joint Cascade Regression Framework for Simultaneous Facial Action Unit Recognition and Facial Landmark Detection

  • Jun 01, 2016
  • Yue Wu +1
  • Conference Article
  • Citations5

A Multi-Genre Urdu Broadcast Speech Recognition System

  • Nov 18, 2021
  • Erbaz Khan +3
  • Conference Article

Confusability Measure Based Lexicon Optimization for Fast LVCSR Decoding

  • Aug 24, 2014
  • Advanced science and technology letters
  • Nam Kyun Kim +2
  • Conference Article
  • Citations4

Minimum Kullback-Leibler distance based multivariate Gaussian feature adaptation for distant-talking speech recognition

  • May 17, 2004
  • Yue Pan +1
  • Book Chapter
  • Citations25

Deep Neural Network Based Continuous Speech Recognition for Serbian Using the Kaldi Toolkit

  • Jan 01, 2015
  • Branislav Popović +4
  • Conference Article
  • Citations15

Noise robust feature extraction based on extended weighted linear prediction in LVCSR

  • Aug 27, 2011
  • Sami Keronen +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.