• Home
  • Search
  • Noise-Robust Speech Recognition System based on Multimodal Audio-Visual Approach Using Different Deep Learning Classification Techniques
  • Cite Icon7
  • https://doi.org/10.21608/ejle.2020.22022.1002Copy DOI Icon

Noise-Robust Speech Recognition System based on Multimodal Audio-Visual Approach Using Different Deep Learning Classification Techniques

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This paper extends an earlier work on designing a speech recognition system based on Hidden Markov Model (HMM) classification technique of using visual modality in addition to audio modality[1]. Improved off traditional HMM-based Automatic Speech Recognition (ASR) accuracy is achieved by implementing a technique using either RNN-based or CNN-based approach. This research is intending to deliver two contributions: The first contribution is the methodology of choosing the visual features by comparing different visual features extraction methods like Discrete Cosine Transform (DCT), blocked DCT, and Histograms of Oriented Gradients with Local Binary Patterns (HOG+LBP), and applying different dimension reduction techniques like Principal Component Analysis (PCA), auto-encoder, Linear Discriminant Analysis (LDA), t-distributed Stochastic Neighbor Embedding (t-SNE) to find the most effective features vector size. Then the obtained visual features are early integrated with the audio features obtained by using Mel Frequency Cepstral Coefficients (MFCCs) and feed the combined audio-visual feature vector to the classification process. The second contribution of this research is the methodology of developing the classification process using deep learning by comparing different Deep Neural Network (DNN) architectures like Bidirectional Long-Short Term Memory (BiLSTM) and Convolution Neural Network (CNN) with the traditional HMM. The proposed model is evaluated on two multi-speakers AV-ASR datasets named AVletters and GRID with different SNR. The model performs speaker-independent experiments in AVlettter dataset and speaker-dependent in GRID dataset.

Loading PDF

Similar Papers
  • Research Article
  • Citations19

Application of Deep Learning for Reservoir Porosity Prediction and Self Organizing Map for Lithofacies Prediction

  • Aug 31, 2024
  • Journal of Applied Geophysics
  • Mazahir Hussain +5
  • PDF
  • Research Article
  • Citations23

A Learning-Based Vehicle-Cloud Collaboration Approach for Joint Estimation of State-of-Energy and State-of-Health

  • Dec 04, 2022
  • Sensors (Basel, Switzerland)
  • Peng Mei +5
  • Research Article

Deep Learning and Traditional Models for Wind Speed Forecasting in Saudi Arabia

  • May 31, 2025
  • ELECTRON Jurnal Ilmiah Teknik Elektro
  • Ikhsan Hidayat +1
  • Research Article
  • Citations221

Multimodal multitask deep learning model for Alzheimer’s disease progression detection based on time series data

  • Jun 01, 2020
  • Neurocomputing
  • Shaker El-Sappagh +3
  • PDF
  • Research Article
  • Citations98

Batteries State of Health Estimation via Efficient Neural Networks With Multiple Channel Charging Profiles

  • Dec 30, 2020
  • IEEE Access
  • Noman Khan +5
  • Research Article
  • Citations5

One-Dimensional Shallow Neural Network Using Non-Fiducial Based Segmented Electrocardiogram for User Identification System

  • Jan 01, 2023
  • IEEE Access
  • Yejin Kim +2
  • Conference Article
  • Citations5

Integration of articulatory knowledge and voicing features based on DNN/HMM for Mandarin speech recognition

  • Jul 01, 2015
  • Ying-Wei Tan +3
  • Conference Article
  • Citations21

Stacked Convolutional Bidirectional LSTM Recurrent Neural Network for Bearing Anomaly Detection in Rotating Machinery Diagnostics

  • Jul 01, 2018
  • Kwangsuk Lee +4
  • Conference Article
  • Citations1

Local and Global Feature Based Hybrid Deep Learning Model for Bangla Parts of Speech Tagging

  • May 21, 2021
  • Muntasir Hoq +2
  • Research Article

Enhanced Image Captioning using Bidirectional Long Short-Term Memory and Convolutional Neural Networks

  • Mar 30, 2024
  • International Journal of Scientific Methods in Engineering and Management
  • Sushma Jaiswal +2
  • PDF
  • Research Article
  • Citations6

A Method for Predicting Tool Remaining Useful Life: Utilizing BiLSTM Optimized by an Enhanced NGO Algorithm

  • Aug 02, 2024
  • Mathematics
  • Jianwei Wu +2
  • Research Article
  • Citations23

Investigation of acoustic and visual features for pig cough classification

  • Jun 01, 2022
  • Biosystems Engineering
  • Nan Ji +7
  • Research Article
  • Citations12

Comparison of Feature Extraction Mel Frequency Cepstral Coefficients and Linear Predictive Coding in Automatic Speech Recognition for Indonesian

  • Mar 01, 2017
  • TELKOMNIKA (Telecommunication Computing Electronics and Control)
  • Sukmawati Nur Endah +2
  • Research Article
  • Citations31

Prediction of PM2.5 concentration in urban agglomeration of China by hybrid network model

  • Sep 12, 2022
  • Journal of Cleaner Production
  • Shuaiwen Wu +1
  • Conference Article
  • Citations5

Gradient Boosting Decision Tree with LSTM for Investment Prediction

  • Apr 23, 2025
  • Chang Yu +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.