TensorX
返回文献探索

Paper · arXiv 2401.02839

Pheme: Efficient and Conversational Speech Generation

Paweł Budzianowski, Taras Sereda, Tomasz Cichy, Ivan Vulić

18 upvotesJanuary 5, 2024arXiv 预印本
AI 摘要

The Pheme model series achieves compact and high-quality voice generation with parallel processing, efficient training on smaller datasets, and improved voice quality through distillation.

speech generationhierarchical neural audio codecsTTS modelsautoregressive naturereal-time usagecompact modelsparallel speech generationnatural conversational speechteacher-student distillationpretrained Pheme checkpoints

Abstract

In recent years, speech generation has seen remarkable progress, now achieving one-shot generation capability that is often virtually indistinguishable from real human voice. Integrating such advancements in speech generation with large language models might revolutionize a wide range of applications. However, certain applications, such as assistive conversational systems, require natural and conversational speech generation tools that also operate efficiently in real time. Current state-of-the-art models like VALL-E and SoundStorm, powered by hierarchical neural audio codecs, require large neural components and extensive training data to work well. In contrast, MQTTS aims to build more compact conversational TTS models while capitalizing on smaller-scale real-life conversational speech data. However, its autoregressive nature yields high inference latency and thus limits its real-time usage. In order to mitigate the current limitations of the state-of-the-art TTS models while capitalizing on their strengths, in this work we introduce the Pheme model series that 1) offers compact yet high-performing models, 2) allows for parallel speech generation of 3) natural conversational speech, and 4) it can be trained efficiently on smaller-scale conversational data, cutting data demands by more than 10x but still matching the quality of the autoregressive TTS models. We also show that through simple teacher-student distillation we can meet significant improvements in voice quality for single-speaker setups on top of pretrained Pheme checkpoints, relying solely on synthetic speech generated by much larger teacher models. Audio samples and pretrained models are available online.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号