TensorX
返回文献探索

Paper · arXiv 2501.10045

HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution

Shengkui Zhao, Kun Zhou, Zexu Pan, Yukun Ma, Chong Zhang, Bin Ma

10 upvotesJanuary 17, 2025arXiv 预印本
AI 摘要

HiFi-SR, a unified GAN-based model, achieves high-fidelity speech super-resolution using a transformer-convolutional generator and multi-scale discriminators, outperforming existing methods across various scenarios.

generative adversarial networksGANsspeech super-resolutionSRmel-spectrogramsunified networkend-to-end adversarial trainingtransformer-convolutional generatorlatent space representationstime-domain waveformsmulti-bandmulti-scale time-frequency discriminatormulti-scale mel-reconstruction loss

Abstract

The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently trained and concatenated networks may lead to inconsistent representations and poor speech quality, especially in out-of-domain scenarios. In this work, we propose HiFi-SR, a unified network that leverages end-to-end adversarial training to achieve high-fidelity speech super-resolution. Our model features a unified transformer-convolutional generator designed to seamlessly handle both the prediction of latent representations and their conversion into time-domain waveforms. The transformer network serves as a powerful encoder, converting low-resolution mel-spectrograms into latent space representations, while the convolutional network upscales these representations into high-resolution waveforms. To enhance high-frequency fidelity, we incorporate a multi-band, multi-scale time-frequency discriminator, along with a multi-scale mel-reconstruction loss in the adversarial training process. HiFi-SR is versatile, capable of upscaling any input speech signal between 4 kHz and 32 kHz to a 48 kHz sampling rate. Experimental results demonstrate that HiFi-SR significantly outperforms existing speech SR methods across both objective metrics and ABX preference tests, for both in-domain and out-of-domain scenarios (https://github.com/modelscope/ClearerVoice-Studio).

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution | TensorX