TensorX
返回文献探索

Paper · arXiv 2401.10032

FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder

Tan Dat Nguyen, Ji-Hoon Kim, Youngjoon Jang, Jaehun Kim, Joon Son Chung

13 upvotesJanuary 18, 2024arXiv 预印本
AI 摘要

FreGrad is a lightweight and fast diffusion-based vocoder that uses discrete wavelet transform for simplified feature representation, frequency-aware dilated convolution for improved frequency accuracy, and optimization techniques to achieve faster training and inference while maintaining audio quality.

diffusion-based vocoderdiscrete wavelet transformfrequency-aware dilated convolutionwaveformsub-band waveletsfrequency awarenessaudio generationmodel sizetraining timeinference speed

Abstract

The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps FreGrad to operate on a simple and concise feature space, (2) We design a frequency-aware dilated convolution that elevates frequency awareness, resulting in generating speech with accurate frequency information, and (3) We introduce a bag of tricks that boosts the generation quality of the proposed model. In our experiments, FreGrad achieves 3.7 times faster training time and 2.2 times faster inference speed compared to our baseline while reducing the model size by 0.6 times (only 1.78M parameters) without sacrificing the output quality. Audio samples are available at: https://mm.kaist.ac.kr/projects/FreGrad.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder | TensorX