TensorX
返回文献探索

Paper · arXiv 2410.16048

Continuous Speech Synthesis using per-token Latent Diffusion

Arnon Turetzky, Nimrod Shabtay, Slava Shechtman, Hagai Aronowitz, David Haws, Ron Hoory, Avihu Dekel

30 upvotesOctober 21, 2024arXiv 预印本
AI 摘要

SALAD, a per-token latent diffusion model using continuous representations, achieves superior intelligibility in zero-shot text-to-speech without compromising speech quality and speaker similarity.

autoregressive transformer modelslatent diffusion modelvariable-length outputsdiffusion headsemantic tokensdiscrete speech synthesis techniquescontinuous speech modeling techniquesintelligibility scorespeech qualityspeaker similarity

Abstract

The success of autoregressive transformer models with discrete tokens has inspired quantization-based approaches for continuous modalities, though these often limit reconstruction quality. We therefore introduce SALAD, a per-token latent diffusion model for zero-shot text-to-speech, that operates on continuous representations. SALAD builds upon the recently proposed expressive diffusion head for image generation, and extends it to generate variable-length outputs. Our approach utilizes semantic tokens for providing contextual information and determining the stopping condition. We suggest three continuous variants for our method, extending popular discrete speech synthesis techniques. Additionally, we implement discrete baselines for each variant and conduct a comparative analysis of discrete versus continuous speech modeling techniques. Our results demonstrate that both continuous and discrete approaches are highly competent, and that SALAD achieves a superior intelligibility score while obtaining speech quality and speaker similarity on par with the ground-truth audio.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Continuous Speech Synthesis using per-token Latent Diffusion | TensorX