TensorX
返回文献探索

Paper · arXiv 2305.07243

Better speech synthesis through scaling

James Betker

6 upvotesMay 12, 2023arXiv 预印本
AI 摘要

TorToise is an expressive multi-voice speech synthesis system that applies autoregressive transformers and DDPMs from image generation to text-to-speech.

autoregressive transformersDDPMs

Abstract

In recent years, the field of image generation has been revolutionized by the application of autoregressive transformers and DDPMs. These approaches model the process of image generation as a step-wise probabilistic processes and leverage large amounts of compute and data to learn the image distribution. This methodology of improving performance need not be confined to images. This paper describes a way to apply advances in the image generative domain to speech synthesis. The result is TorToise -- an expressive, multi-voice text-to-speech system. All model code and trained weights have been open-sourced at https://github.com/neonbjb/tortoise-tts.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Better speech synthesis through scaling | TensorX