TensorX
返回文献探索

Paper · arXiv 2306.02982

PolyVoice: Language Models for Speech to Speech Translation

Qianqian Dong, Zhiying Huang, Chen Xu, Yunlong Zhao, Kexin Wang, Xuxin Cheng, Tom Ko, Qiao Tian, Tang Li, Fengpeng Yue, Ye Bai, Xi Chen, Lu Lu, Zejun Ma, Yuping Wang, Mingxuan Wang, Yuxuan Wang

4 upvotesJune 5, 2023arXiv 预印本
AI 摘要

PolyVoice, a language model-based framework for speech-to-speech translation, uses discretized speech units and VALL-E X for high-quality voice and style preservation in translation.

language model-based frameworkspeech-to-speech translationtranslation language modelspeech synthesis language modeldiscretized speech unitsunsupervised generationVALL-E Xunit-based audio language model

Abstract

We propose PolyVoice, a language model-based framework for speech-to-speech translation (S2ST) system. Our framework consists of two language models: a translation language model and a speech synthesis language model. We use discretized speech units, which are generated in a fully unsupervised way, and thus our framework can be used for unwritten languages. For the speech synthesis part, we adopt the existing VALL-E X approach and build a unit-based audio language model. This grants our framework the ability to preserve the voice characteristics and the speaking style of the original speech. We examine our system on Chinese rightarrow English and English rightarrow Spanish pairs. Experimental results show that our system can generate speech with high translation quality and audio quality. Speech samples are available at https://speechtranslation.github.io/polyvoice.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
PolyVoice: Language Models for Speech to Speech Translation | TensorX