• Home
  • Search
  • UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
  • Cite Icon2
  • https://doi.org/10.1109/icassp49357.2023.10096194Copy DOI Icon

UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction

  • Jun 4, 2023
  • Jiaxin Guo +11 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Error correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works usually adopt end-to-end models and has strong dependency on Pseudo Paired Data and Original Paired Data. But when only pre-training on Pseudo Paired Data, previous models have negative effect on correction. While fine-tuning on Original Paired Data, the source side data must be transcribed by a well-trained ASR model, which takes a lot of time and not universal. In this paper, we propose UCorrect, an unsupervised Detector-Generator-Selector framework for ASR Error Correction. UCorrect has no dependency on the training data mentioned before. The whole procedure is first to detect whether the character is erroneous, then to generate some candidate characters and finally to select the most confident one to replace the error character. Experiments on the public AISHELL-1 dataset and WenetSpeech dataset show the effectiveness of UCorrect for ASR error correction: 1) it achieves significant WER reduction, achieves 6.83% even without fine-tuning and 14.29% after fine-tuning; 2) it outperforms the popular NAR correction models by a large margin with a competitive low latency; and 3) it is an universal method, as it reduces all WERs of the ASR model with different decoding strategies and reduces all WERs of ASR models trained on different scale datasets.

Similar Papers
  • Research Article

Enhancing Armenian Automatic Speech Recognition Performance: A Comprehensive Strategy for Speed, Accuracy, and Linguistic Refinement

  • Jan 01, 2024
  • International Journal Of Scientific Advances
  • Varuzhan H. Baghdasaryan
  • Conference Article
  • Citations5

Bangla-Wave: Improving Bangla Automatic Speech Recognition Utilizing N-gram Language Models

  • Feb 23, 2023
  • Mohammed Rakib +3
  • Research Article

Confidence Gated Fusion: Dynamic Language Model Integration for Adapting Pretrained Multilingual ASR Models with Text-Only Data

  • Jan 01, 2026
  • Procedia Computer Science
  • Nader Essam +4
  • Research Article
  • Citations25

Enhancements in automatic Kannada speech recognition system by background noise elimination and alternate acoustic modelling

  • Jan 22, 2020
  • International Journal of Speech Technology
  • G Thimmaraja Yadava +1
  • Research Article

Running Automatic Speech Recognition (ASR) Model in the Context of Real Time Data Streaming Architecture

  • Jul 01, 2025
  • Proceedings of the International Conference on Business Excellence
  • Robert Cristian Necula +1
  • Preprint Article

Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs

  • Feb 25, 2026
  • arXiv (Cornell University)
  • Pietroń, Marcin +8
  • Conference Article
  • Citations4

Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models

  • Jan 01, 2025
  • Hanin Atwany +4
  • Conference Article
  • Citations1

An End to End Spoken Dialogue System to Access the Agricultural Commodity Price Information in Kannada Language/Dialects

  • Dec 01, 2017
  • G Thimmaraja Yadava +2
  • Conference Article
  • Citations10

Predicting Word Error Rate for Reverberant Speech

  • May 01, 2020
  • Hannes Gamper +3
  • Research Article

HiACC: Hinglish adult & children code-switched corpus

  • Jul 17, 2025
  • Data in Brief
  • Shruti Singh +2
  • Conference Article
  • Citations2

Retrieving phrases by selecting the history: application to automatic speech recognition

  • Sep 16, 2002
  • David Langlois +2
  • Conference Article
  • Citations8

Omni-Sparsity DNN: Fast Sparsity Optimization for On-Device Streaming E2E ASR Via Supernet

  • May 23, 2022
  • Haichuan Yang +8
  • Conference Article

VoCare AI: A Multi-Agent LLM Workflow for Improved Clinic Operational Efficiency

  • Oct 27, 2025
  • Derrick Lim Jo Han +2
  • Conference Article
  • Citations25

Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement

  • Aug 30, 2021
  • Sefik Emre Eskimez +7
  • Conference Article
  • Citations23

Language Model Estimation for Optimizing End-to-end Performance of a Natural Language Call Routing System

  • Mar 18, 2005
  • V Goel +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.