TensorX
返回文献探索

Paper · arXiv 2309.07314

AudioSR: Versatile Audio Super-resolution at Scale

Haohe Liu, Ke Chen, Qiao Tian, Wenwu Wang, Mark D. Plumbley

29 upvotesSeptember 13, 2023arXiv 预印本
AI 摘要

A diffusion-based generative model, AudioSR, achieves robust audio super-resolution across various audio types and bandwidths, enhancing generation quality for different audio models.

diffusion-based generative modelaudio super-resolutionaudio typessound effectsmusicspeechbandwidth rangehigh-resolution audiosampling rateAudioLDMFastspeech2MusicGen

Abstract

Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications. Previous methods have limitations such as the limited scope of audio types (e.g., music, speech) and specific bandwidth settings they can handle (e.g., 4kHz to 8kHz). In this paper, we introduce a diffusion-based generative model, AudioSR, that is capable of performing robust audio super-resolution on versatile audio types, including sound effects, music, and speech. Specifically, AudioSR can upsample any input audio signal within the bandwidth range of 2kHz to 16kHz to a high-resolution audio signal at 24kHz bandwidth with a sampling rate of 48kHz. Extensive objective evaluation on various audio super-resolution benchmarks demonstrates the strong result achieved by the proposed model. In addition, our subjective evaluation shows that AudioSR can acts as a plug-and-play module to enhance the generation quality of a wide range of audio generative models, including AudioLDM, Fastspeech2, and MusicGen. Our code and demo are available at https://audioldm.github.io/audiosr.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
AudioSR: Versatile Audio Super-resolution at Scale | TensorX