TensorX
返回文献探索

Paper · arXiv 2401.09084

UniVG: Towards UNIfied-modal Video Generation

Ludan Ruan, Lei Tian, Chuanwei Huang, Xu Zhang, Xinyan Xiao

17 upvotesJanuary 17, 2024arXiv 预印本
AI 摘要

The Unified-modal Video Generation system uses Multi-condition Cross Attention and Biased Gaussian Noise to handle diverse video generation tasks across text and image modalities, achieving superior performance on benchmarks.

diffusion based video generationUnified-modal Video Generationgenerative freedomMulti-condition Cross AttentionBiased Gaussian NoiseFr\'echet Video Distance (FVD)MSR-VTTGen2

Abstract

Diffusion based video generation has received extensive attention and achieved considerable success within both the academic and industrial communities. However, current efforts are mainly concentrated on single-objective or single-task video generation, such as generation driven by text, by image, or by a combination of text and image. This cannot fully meet the needs of real-world application scenarios, as users are likely to input images and text conditions in a flexible manner, either individually or in combination. To address this, we propose a Unified-modal Video Genearation system that is capable of handling multiple video generation tasks across text and image modalities. To this end, we revisit the various video generation tasks within our system from the perspective of generative freedom, and classify them into high-freedom and low-freedom video generation categories. For high-freedom video generation, we employ Multi-condition Cross Attention to generate videos that align with the semantics of the input images or text. For low-freedom video generation, we introduce Biased Gaussian Noise to replace the pure random Gaussian Noise, which helps to better preserve the content of the input conditions. Our method achieves the lowest Fr\'echet Video Distance (FVD) on the public academic benchmark MSR-VTT, surpasses the current open-source methods in human evaluations, and is on par with the current close-source method Gen2. For more samples, visit https://univg-baidu.github.io.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
UniVG: Towards UNIfied-modal Video Generation | TensorX