TensorX
返回文献探索

Paper · arXiv 2307.04686

VampNet: Music Generation via Masked Acoustic Token Modeling

Hugo Flores Garcia, Prem Seetharaman, Rithesh Kumar, Bryan Pardo

22 upvotesJuly 10, 2023arXiv 预印本
AI 摘要

VampNet, a non-autoregressive transformer-based model, synthesizes and manipulates music with high-fidelity using variable masking and prompting techniques.

masked acoustic token modelingmusic synthesismusic compressionmusic inpaintingbidirectional transformer architecturepromptinghigh-fidelity musical waveformsmusic co-creation

Abstract

We introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by applying a variety of masking approaches (called prompts) during inference. VampNet is non-autoregressive, leveraging a bidirectional transformer architecture that attends to all tokens in a forward pass. With just 36 sampling passes, VampNet can generate coherent high-fidelity musical waveforms. We show that by prompting VampNet in various ways, we can apply it to tasks like music compression, inpainting, outpainting, continuation, and looping with variation (vamping). Appropriately prompted, VampNet is capable of maintaining style, genre, instrumentation, and other high-level aspects of the music. This flexible prompting capability makes VampNet a powerful music co-creation tool. Code and audio samples are available online.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VampNet: Music Generation via Masked Acoustic Token Modeling | TensorX