TensorX
返回文献探索

Paper · arXiv 2407.02869

PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation

Zeyu Xie, Xuenan Xu, Zhizheng Wu, Mengyue Wu

19 upvotesJuly 3, 2024arXiv 预印本
AI 摘要

PicoAudio, a temporal-controlled audio generation framework, enhances timestamp and occurrence frequency controllability through tailored model design and fine-grained audio-text data.

temporal controllabilityaudio generation frameworktemporal informationdata crawlingsegmentationfilteringsimulationfine-grained temporally-aligned audio-text data

Abstract

Recently, audio generation tasks have attracted considerable research interests. Precise temporal controllability is essential to integrate audio generation with real applications. In this work, we propose a temporal controlled audio generation framework, PicoAudio. PicoAudio integrates temporal information to guide audio generation through tailored model design. It leverages data crawling, segmentation, filtering, and simulation of fine-grained temporally-aligned audio-text data. Both subjective and objective evaluations demonstrate that PicoAudio dramantically surpasses current state-of-the-art generation models in terms of timestamp and occurrence frequency controllability. The generated samples are available on the demo website https://PicoAudio.github.io.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation | TensorX