TensorX
返回文献探索

Paper · arXiv 2409.00587

FLUX that Plays Music

Zhengcong Fei, Mingyuan Fan, Changqian Yu, Junshi Huang

33 upvotesSeptember 1, 2024arXiv 预印本
AI 摘要

FluxMusic extends diffusion-based rectified flow Transformers to generate music from text by transforming it into a latent VAE space of mel-spectra and using text encoders and modulations for enhanced semantic capture.

diffusion-based rectified flowTransformersFluxlatent VAE spacemel-spectrumattentiondenoised patch predictiontext encodersmodulation mechanismautomatic metricshuman preference evaluations

Abstract

This paper explores a simple extension of diffusion-based rectified flow Transformers for text-to-music generation, termed as FluxMusic. Generally, along with design in advanced Fluxhttps://github.com/black-forest-labs/flux model, we transfers it into a latent VAE space of mel-spectrum. It involves first applying a sequence of independent attention to the double text-music stream, followed by a stacked single music stream for denoised patch prediction. We employ multiple pre-trained text encoders to sufficiently capture caption semantic information as well as inference flexibility. In between, coarse textual information, in conjunction with time step embeddings, is utilized in a modulation mechanism, while fine-grained textual details are concatenated with the music patch sequence as inputs. Through an in-depth study, we demonstrate that rectified flow training with an optimized architecture significantly outperforms established diffusion methods for the text-to-music task, as evidenced by various automatic metrics and human preference evaluations. Our experimental data, code, and model weights are made publicly available at: https://github.com/feizc/FluxMusic.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号