TensorX
返回文献探索

Paper · arXiv 2408.11915

Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound

Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam

8 upvotesAugust 21, 2024arXiv 预印本
AI 摘要

Video-Foley achieves top performance in video-to-sound generation by using RMS for temporal control and semantic timbre prompts without human annotation.

RMSRoot Mean Squareintensity envelopetemporal event conditionsemantic timbre promptsannotation-freeself-supervised learningVideo2RMSRMS2SoundRMS discretizationRMS-ControlNettext-to-audio modelaudio-visual alignment

Abstract

Foley sound synthesis is crucial for multimedia production, enhancing user experience by synchronizing audio and video both temporally and semantically. Recent studies on automating this labor-intensive process through video-to-sound generation face significant challenges. Systems lacking explicit temporal features suffer from poor controllability and alignment, while timestamp-based models require costly and subjective human annotation. We propose Video-Foley, a video-to-sound system using Root Mean Square (RMS) as a temporal event condition with semantic timbre prompts (audio or text). RMS, a frame-level intensity envelope feature closely related to audio semantics, ensures high controllability and synchronization. The annotation-free self-supervised learning framework consists of two stages, Video2RMS and RMS2Sound, incorporating novel ideas including RMS discretization and RMS-ControlNet with a pretrained text-to-audio model. Our extensive evaluation shows that Video-Foley achieves state-of-the-art performance in audio-visual alignment and controllability for sound timing, intensity, timbre, and nuance. Code, model weights, and demonstrations are available on the accompanying website. (https://jnwnlee.github.io/video-foley-demo)

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound | TensorX