• Home
  • Search
  • Voice Command Accuracy Improvement Using Connectionist Temporal Classification in Speech-to-Text Interfaces
  • https://doi.org/10.1109/aistemedu67077.2025.11403879Copy DOI Icon

Voice Command Accuracy Improvement Using Connectionist Temporal Classification in Speech-to-Text Interfaces

  • Dec 1, 2025
  • Shokhista Makhmaraimova +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Voice interfaces are an essential part of contemporary human-computer interaction and speech-to-text (STT) systems allow one to work hands-free and enhance the accessibility. Nevertheless, the precision of voice command recognition is a major issue that is yet to be addressed particularly in noisy settings or in language with complex linguistic features. Current approaches can have issues with the consistency of real-time transcription and are not contextually aware and cannot work with limited or agglutinative languages like Hindi. In order to overcome these shortcomings, Enhancing Voice Command with Connectionist Temporal Classification (EVC-CTC) is proposed in this study. It is a strong framework which makes use of the CTC to enhance sequence matching between speech input and textual output. EVC-CTC combines both context modeling and language specific features to increase recognition accuracy particularly in the presence of real-world variability. The suggested approach is deployed as a modular and scalable STT pipeline that facilitates voice assistants that run on web-based platforms. Continuous voice streams are processed with EVC-CTC to preserve continuity of the context and dynamically respond to the background noise and variations in speakers. Experimental analysis indicates that EVC-CTC is much better than baseline models in Word Error Rate (WER) and semantic accuracy, especially in Hindi voice commands.

Similar Papers
  • Research Article
  • Citations36

Advancing Acoustic-to-Word CTC Model With Attention and Mixed-Units

  • Sep 04, 2019
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Amit Das +4
  • Conference Article
  • Citations67

Optimizing Expected Word Error Rate via Sampling for Speech Recognition

  • Aug 20, 2017
  • Matt Shannon
  • Book Chapter
  • Citations4

Gujarati Language Automatic Speech Recognition Using Integrated Feature Extraction and Hybrid Acoustic Model

  • Jan 01, 2023
  • Mohit Dua +1
  • Conference Article
  • Citations25

Joint Masked CPC And CTC Training For ASR

  • Jun 06, 2021
  • Chaitanya Talnikar +3
  • Conference Article
  • Citations7

Investigating Sequence-Level Normalisation For CTC-Like End-to-End ASR

  • May 23, 2022
  • Zeyu Zhao +1
  • Research Article
  • Citations15

Improving Deep Learning based Automatic Speech Recognition for Gujarati

  • Dec 13, 2021
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Deepang Raval +3
  • Conference Article
  • Citations41

Transliteration Based Approaches to Improve Code-Switched Speech Recognition Performance

  • Dec 01, 2018
  • Jesse Emond +4
  • Research Article
  • Citations4

End-to-end recognition of streaming Japanese speech using CTC and local attention

  • Jan 01, 2020
  • APSIPA Transactions on Signal and Information Processing
  • Jiahao Chen +2
  • Conference Article
  • Citations24

Advancing Multi-Accented Lstm-CTC Speech Recognition Using a Domain Specific Student-Teacher Learning Paradigm

  • Dec 01, 2018
  • Shahram Ghorbani +2
  • Conference Article
  • Citations2

Two-Stage Augmentation and Adaptive CTC Fusion for Improved Robustness of Multi-Stream end-to-end ASR

  • Jan 19, 2021
  • Ruizhi Li +2
  • PDF
  • Research Article
  • Citations21

Accented Speech Recognition Based on End-to-End Domain Adversarial Training of Neural Networks

  • Sep 10, 2021
  • Applied Sciences
  • Hyeong-Ju Na +1
  • Research Article
  • Citations22

"Mm-hm," "Uh-uh": are non-lexical conversational sounds deal breakers for the ambient clinical documentation technology?

  • Jan 23, 2023
  • Journal of the American Medical Informatics Association
  • Brian D Tran +6
  • Research Article
  • Citations25

Enhancements in automatic Kannada speech recognition system by background noise elimination and alternate acoustic modelling

  • Jan 22, 2020
  • International Journal of Speech Technology
  • G Thimmaraja Yadava +1
  • Conference Article
  • Citations1

End-to-end speech recognition with Alignment RNN-Transducer

  • Jul 18, 2021
  • Ying Tian +5
  • Research Article
  • Citations3

CNN-Based Speaker Verification and Speech Recognition in Tibetan

  • Dec 01, 2020
  • Journal of Physics: Conference Series
  • Zhenye Gan +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.