TensorX
返回文献探索

Paper · arXiv 2406.02897

LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Trung Dang, David Aponte, Dung Tran, Kazuhito Koishida

15 upvotesJune 5, 2024arXiv 预印本
AI 摘要

A novel autoregressive language model approach, LiveSpeech, enables low-latency text-to-speech by optimizing token prediction and parallel codebook processing.

generative language modelaudio tokensneural audio codecadaptive codebook loss weightsautoregressive language modellow-latency text-to-speechcodebooksparallel processingcontent accuracyspeaker similarityaudio qualityinference speed

Abstract

Prior works have demonstrated zero-shot text-to-speech by using a generative language model on audio tokens obtained via a neural audio codec. It is still challenging, however, to adapt them to low-latency scenarios. In this paper, we present LiveSpeech - a fully autoregressive language model-based approach for zero-shot text-to-speech, enabling low-latency streaming of the output audio. To allow multiple token prediction within a single decoding step, we propose (1) using adaptive codebook loss weights that consider codebook contribution in each frame and focus on hard instances, and (2) grouping codebooks and processing groups in parallel. Experiments show our proposed models achieve competitive results to state-of-the-art baselines in terms of content accuracy, speaker similarity, audio quality, and inference speed while being suitable for low-latency streaming applications.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes | TensorX