TensorX
返回文献探索

Paper · arXiv 2311.08667

EDMSound: Spectrogram Based Diffusion Models for Efficient and High-Quality Audio Synthesis

Ge Zhu, Yutong Wen, Marc-André Carbonneau, Zhiyao Duan

18 upvotesNovember 15, 2023arXiv 预印本
AI 摘要

EDMSound, an audio diffusion model in the spectrogram domain using an efficient deterministic sampler, achieves state-of-the-art performance with fewer steps and highlights concerns about perceptual similarity in diffusion-based audio generation.

diffusion modelslatent domaincascaded phase recovery modulesspectrogram domainelucidated diffusion modelsEDMefficient deterministic samplerFréchet audio distanceFADDCASE2023 foley sound generation benchmark

Abstract

Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. This poses challenges when generating high-fidelity audio. In this paper, we propose EDMSound, a diffusion-based generative model in spectrogram domain under the framework of elucidated diffusion models (EDM). Combining with efficient deterministic sampler, we achieved similar Fr\'echet audio distance (FAD) score as top-ranked baseline with only 10 steps and reached state-of-the-art performance with 50 steps on the DCASE2023 foley sound generation benchmark. We also revealed a potential concern regarding diffusion based audio generation models that they tend to generate samples with high perceptual similarity to the data from training data. Project page: https://agentcooper2002.github.io/EDMSound/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
EDMSound: Spectrogram Based Diffusion Models for Efficient and High-Quality Audio Synthesis | TensorX