• Home
  • Search
  • Multimodal speech recognition: increasing accuracy using high speed video data
  • Cite Icon26
  • https://doi.org/10.1007/s12193-018-0267-1Copy DOI Icon

Multimodal speech recognition: increasing accuracy using high speed video data

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

To date, multimodal speech recognition systems based on the processing of audio and video signals show significantly better results than their unimodal counterparts. In general, researchers divide the solution of the audio–visual speech recognition problem into two parts. First, in extracting the most informative features from each modality and second, in the most successful way of fusion both modalities. Ultimately, this leads to an improvement in the accuracy of speech recognition. Almost all modern studies use this approach with video data of a standard recording speed of 25 frames per second. The choice of such a recording speed is easily explained, since the vast majority of existing audio–visual databases are recorded with this rate. However, it should be noticed that the number of 25 frames per second is a world standard for many areas and has never been specifically calculated for speech recognition tasks. The main purpose of this study is to investigate the effect brought by the high-speed video data (up to 200 frames per second) on the speech recognition accuracy. And also to find out whether the use of a high-speed video camera makes the speech recognition systems more robust to acoustical noise. To this end, we recorded a database of audio–visual Russian speech with high-speed video recordings, which consists of records of 20 speakers, each of them pronouncing 200 phrases of continuous Russian speech. Experiments performed on this database showed an improvement in the absolute speech recognition rate up to 3.10%. We also proved that the use of the high-speed camera with 200 fps allows achieving better recognition results under different acoustically noisy conditions (signal-to-noise ratio varied between 40 and 0 dB) with different types of noise (e.g. white noise, babble noise).

Similar Papers
  • Research Article
  • Citations25

Combined speech enhancement and auditory modelling for robust distributed speech recognition

  • May 20, 2008
  • Speech Communication
  • Ronan Flynn +1
  • Research Article
  • Citations51

Speech Recognition and In-Vehicle Telematics Devices: Potential Reductions in Driver Distraction

  • Jan 01, 2004
  • International Journal of Speech Technology
  • Marvin C Mccallum +4
  • Research Article
  • Citations13

A comparison of automatic and human speech recognition in null grammar

  • Feb 17, 2012
  • The Journal of the Acoustical Society of America
  • Amit Juneja
  • Research Article
  • Citations35

Audio–visual speech recognition based on regulated transformer and spatio–temporal fusion strategy for driver assistive systems

  • May 09, 2024
  • Expert Systems With Applications
  • Dmitry Ryumin +5
  • Research Article
  • Citations25

Enhancements in automatic Kannada speech recognition system by background noise elimination and alternate acoustic modelling

  • Jan 22, 2020
  • International Journal of Speech Technology
  • G Thimmaraja Yadava +1
  • Research Article
  • Citations7

FEATURE EXTRACTION ALGORITHM USING NEW CEPSTRAL TECHNIQUES FOR ROBUST SPEECH RECOGNITION

  • Apr 24, 2020
  • Malaysian Journal of Computer Science
  • Mohamed Cherif Amara Korba +2
  • Research Article
  • Citations3

Challenges of Automatic Speech Recognition for medical interviews - research for Polish language

  • Jan 01, 2023
  • Procedia Computer Science
  • Karolina Kuligowska +2
  • Book Chapter
  • Citations3

Phonemes: An Explanatory Study Applied to Identify a Speaker

  • Jan 01, 2020
  • Saritha Kinkiri +2
  • Conference Article
  • Citations3

An End-to-end Speech Recognition Algorithm based on Attention Mechanism

  • Jul 01, 2020
  • Jia-Nan Chen +5
  • Conference Article
  • Citations6

A New Corpus of Elderly Japanese Speech for Acoustic Modeling, and a Preliminary Investigation of Dialect-Dependent Speech Recognition

  • Oct 01, 2019
  • Meiko Fukuda +4
  • Book Chapter

Error Analysis and Improving the Speech Recognition Accuracy on Telugu Language

  • Jan 01, 2012
  • Lecture notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering
  • N Usha Rani +1
  • Conference Article
  • Citations4

End-to-end Visual Speech Recognition for Human-Robot Interaction

  • Jan 01, 2022
  • Denis Ivanko +2
  • Research Article
  • Citations1

Raw acoustic-articulatory multimodal dysarthric speech recognition

  • Jan 01, 2026
  • Computer Speech & Language
  • Zhengjun Yue +4
  • PDF
  • Research Article
  • Citations21

Accented Speech Recognition Based on End-to-End Domain Adversarial Training of Neural Networks

  • Sep 10, 2021
  • Applied Sciences
  • Hyeong-Ju Na +1
  • Research Article
  • Citations2

Nonlinear filtering for recognition of phase-encoded images

  • Mar 10, 1998
  • Applied Optics
  • Bahram Javidi +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.