TensorX
返回文献探索

Paper · arXiv 2409.09214

Seed-Music: A Unified Framework for High Quality and Controlled Music Generation

Ye Bai, Haonan Chen, Jitong Chen, Zhuo Chen, Yi Deng, Xiaohong Dong, Lamtharn Hantrakul, Weituo Hao, Qingqing Huang, Zhongyi Huang, Dongya Jia, Feihu La, Duc Le, Bochen Li, Chumin Li, Hui Li, Xingxing Li, Shouda Liu, Wei-Tsung Lu, Yiqing Lu, Andrew Shaw, Janne Spijkervet, Yakun Sun, Bo Wang, Ju-Chiang Wang, Yuping Wang, Yuxuan Wang, Ling Xu, Yifeng Yang, Chao Yao, Shuo Zhang, Yang Zhang, Yilin Zhang, Hang Zhao, Ziyi Zhao, Dejian Zhong, Shicen Zhou, Pei Zou

54 upvotesSeptember 13, 2024arXiv 预印本
AI 摘要

Seed-Music generates high-quality music with style control using a unified framework combining auto-regressive language modeling and diffusion approaches, supporting vocal music generation and post-production editing.

auto-regressive language modelingdiffusion approachescontrolled music generationvocal music generationmulti-modal inputsstyle descriptionsaudio referencesmusical scoresvoice promptspost-production editinglyrics editingvocal melodies editing

Abstract

We introduce Seed-Music, a suite of music generation systems capable of producing high-quality music with fine-grained style control. Our unified framework leverages both auto-regressive language modeling and diffusion approaches to support two key music creation workflows: controlled music generation and post-production editing. For controlled music generation, our system enables vocal music generation with performance controls from multi-modal inputs, including style descriptions, audio references, musical scores, and voice prompts. For post-production editing, it offers interactive tools for editing lyrics and vocal melodies directly in the generated audio. We encourage readers to listen to demo audio examples at https://team.doubao.com/seed-music .

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Seed-Music: A Unified Framework for High Quality and Controlled Music Generation | TensorX