TensorX
返回文献探索

Paper · arXiv 2405.15216

Denoising LM: Pushing the Limits of Error Correction Models for Speech Recognition

Zijin Gu, Tatiana Likhomanenko, He Bai, Erik McDermott, Ronan Collobert, Navdeep Jaitly

15 upvotesMay 24, 2024arXiv 预印本
AI 摘要

Denoising language models (DLMs) trained with synthetic data achieve state-of-the-art ASR performance on Librispeech without external audio data, surpassing conventional language models and some self-supervised methods.

Denoising LMscaled error correction modeltext-to-speech systemsASR errorssynthetic datamulti-speaker TTS systemsnoise augmentation strategiesTransformer-CTC ASRword error rate (WER)test-cleantest-otherLibrispeechbeam-search rescoring

Abstract

Language models (LMs) have long been used to improve results of automatic speech recognition (ASR) systems, but they are unaware of the errors that ASR systems make. Error correction models are designed to fix ASR errors, however, they showed little improvement over traditional LMs mainly due to the lack of supervised training data. In this paper, we present Denoising LM (DLM), which is a scaled error correction model trained with vast amounts of synthetic data, significantly exceeding prior attempts meanwhile achieving new state-of-the-art ASR performance. We use text-to-speech (TTS) systems to synthesize audio, which is fed into an ASR system to produce noisy hypotheses, which are then paired with the original texts to train the DLM. DLM has several key ingredients: (i) up-scaled model and data; (ii) usage of multi-speaker TTS systems; (iii) combination of multiple noise augmentation strategies; and (iv) new decoding techniques. With a Transformer-CTC ASR, DLM achieves 1.5% word error rate (WER) on test-clean and 3.3% WER on test-other on Librispeech, which to our knowledge are the best reported numbers in the setting where no external audio data are used and even match self-supervised methods which use external audio data. Furthermore, a single DLM is applicable to different ASRs, and greatly surpassing the performance of conventional LM based beam-search rescoring. These results indicate that properly investigated error correction models have the potential to replace conventional LMs, holding the key to a new level of accuracy in ASR systems.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Denoising LM: Pushing the Limits of Error Correction Models for Speech Recognition | TensorX