TensorX
返回文献探索

Paper · arXiv 2406.04324

SF-V: Single Forward Video Generation Model

Zhixing Zhang, Yanyu Li, Yushu Wu, Yanwu Xu, Anil Kag, Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Junli Cao, Dimitris Metaxas, Sergey Tulyakov, Jian Ren

24 upvotesJune 6, 2024arXiv 预印本
AI 摘要

Adversarial training transforms a multi-step diffusion video generation model into a single-step high-quality video synthesis model with reduced computational cost.

diffusion-based video generationiterative denoisingadversarial trainingpre-trained video diffusion modelssingle-step video generationStable Video Diffusiontemporal dependenciesspatial dependenciescompetitive generation qualityreal-time video synthesisreal-time video editing

Abstract

Diffusion-based video generation models have demonstrated remarkable success in obtaining high-fidelity videos through the iterative denoising process. However, these models require multiple denoising steps during sampling, resulting in high computational costs. In this work, we propose a novel approach to obtain single-step video generation models by leveraging adversarial training to fine-tune pre-trained video diffusion models. We show that, through the adversarial training, the multi-steps video diffusion model, i.e., Stable Video Diffusion (SVD), can be trained to perform single forward pass to synthesize high-quality videos, capturing both temporal and spatial dependencies in the video data. Extensive experiments demonstrate that our method achieves competitive generation quality of synthesized videos with significantly reduced computational overhead for the denoising process (i.e., around 23times speedup compared with SVD and 6times speedup compared with existing works, with even better generation quality), paving the way for real-time video synthesis and editing. More visualization results are made publicly available at https://snap-research.github.io/SF-V.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号