- https://doi.org/10.1109/i-pact65952.2025.11307948
Speech Recognition Using Deep Learning Techniques: A Comparative Study
- Sep 25, 2025
- Hasini E +5 more
Automatic Speech Recognition (ASR) systems have undergone tremendous advancement over the last few years but remain challenged with errors in transcription and readability, particularly in spontaneous speech. In this paper, a hybrid ASR system that employs the T5 transformer model in postediting is introduced for improving transcription quality and coherence. Our methodology focuses on utilizing a conventional ASR pipeline to obtain preliminary transcripts followed by a fine-tuned version of the T5 model, which refines errors and makes the transcript readable. We experiment with the system on several corpora, producing significant gains in word error rate (WER) and linguistic clarity. Experimental metrics indicate that the hybrid method not only surpasses baseline ASR models in precision but also achieves greater readability, rendering it especially appropriate for realistic applications like subtitles, meeting records, and virtual assistants. Further, we also investigate the versatility of the model to accommodate variability in speaking behavior and domain terms. The system we suggest points towards the capabilities of transformer-based post-processing in transforming ASR output to make it more human-like and contextually correct. Our research also discusses the computational efficiency of the process, which ensures its practicality for real-time deployment. We also test the robustness of the model over various conditions of noise and speech accent and prove its generalizability.