TensorX
返回文献探索

Paper · arXiv 2407.14358

Stable Audio Open

Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons

29 upvotesJuly 19, 2024arXiv 预印本
AI 摘要

An open-access text-to-audio model trained with Creative Commons data achieves competitive performance, particularly in high-quality stereo sound synthesis.

open generative modelsfine-tunestext-to-audio modelsCreative Commons dataopen-weightsFDopenl3stereo sound synthesis

Abstract

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Stable Audio Open | TensorX