TensorX
返回文献探索

Paper · arXiv 2406.18009

E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker, Canrun Li, Chung-Hsien Tsai, Zhen Xiao, Hemin Yang, Zirun Zhu, Min Tang, Xu Tan, Yanqing Liu, Sheng Zhao, Naoyuki Kanda

22 upvotesJune 26, 2024arXiv 预印本
AI 摘要

E2 TTS is a non-autoregressive zero-shot text-to-speech system using flow-matching-based mel spectrogram generation without additional components, achieving state-of-the-art performance.

Embarrassingly Easy Text-to-SpeechE2 TTSfully non-autoregressivezero-shot text-to-speechflow-matching-basedmel spectrogramaudio infillingduration modelgrapheme-to-phonememonotonic alignment searchVoiceboxNaturalSpeech 3

Abstract

This paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-art speaker similarity and intelligibility. In the E2 TTS framework, the text input is converted into a character sequence with filler tokens. The flow-matching-based mel spectrogram generator is then trained based on the audio infilling task. Unlike many previous works, it does not require additional components (e.g., duration model, grapheme-to-phoneme) or complex techniques (e.g., monotonic alignment search). Despite its simplicity, E2 TTS achieves state-of-the-art zero-shot TTS capabilities that are comparable to or surpass previous works, including Voicebox and NaturalSpeech 3. The simplicity of E2 TTS also allows for flexibility in the input representation. We propose several variants of E2 TTS to improve usability during inference. See https://aka.ms/e2tts/ for demo samples.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS | TensorX