TensorX
返回文献探索

Paper · arXiv 2401.01792

CoMoSVC: Consistency Model-based Singing Voice Conversion

Yiwen Lu, Zhen Ye, Wei Xue, Xu Tan, Qifeng Liu, Yike Guo

10 upvotesJanuary 3, 2024arXiv 预印本
AI 摘要

CoMoSVC is a consistency model-based singing voice conversion method that achieves fast inference speed and high-quality audio generation.

diffusion-based SVCconsistency modeldiffusion-based teacher modelstudent modelself-consistency propertiesone-step sampling

Abstract

The diffusion-based Singing Voice Conversion (SVC) methods have achieved remarkable performances, producing natural audios with high similarity to the target timbre. However, the iterative sampling process results in slow inference speed, and acceleration thus becomes crucial. In this paper, we propose CoMoSVC, a consistency model-based SVC method, which aims to achieve both high-quality generation and high-speed sampling. A diffusion-based teacher model is first specially designed for SVC, and a student model is further distilled under self-consistency properties to achieve one-step sampling. Experiments on a single NVIDIA GTX4090 GPU reveal that although CoMoSVC has a significantly faster inference speed than the state-of-the-art (SOTA) diffusion-based SVC system, it still achieves comparable or superior conversion performance based on both subjective and objective metrics. Audio samples and codes are available at https://comosvc.github.io/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
CoMoSVC: Consistency Model-based Singing Voice Conversion | TensorX