TensorX
返回文献探索

Paper · arXiv 2407.07667

VEnhancer: Generative Space-Time Enhancement for Video Generation

Jingwen He, Tianfan Xue, Dongyang Liu, Xinqi Lin, Peng Gao, Dahua Lin, Yu Qiao, Wanli Ouyang, Ziwei Liu

16 upvotesJuly 10, 2024arXiv 预印本
AI 摘要

VEnhancer enhances video quality by increasing spatial and temporal resolution using a unified video diffusion model and video ControlNet with space-time data augmentation and video-aware conditioning.

generative space-time enhancementvideo diffusion modelvideo ControlNetspace-time data augmentationvideo-aware conditioningtext-to-videovideo super-resolutionspace-time super-resolutionVBench

Abstract

We present VEnhancer, a generative space-time enhancement framework that improves the existing text-to-video results by adding more details in spatial domain and synthetic detailed motion in temporal domain. Given a generated low-quality video, our approach can increase its spatial and temporal resolution simultaneously with arbitrary up-sampling space and time scales through a unified video diffusion model. Furthermore, VEnhancer effectively removes generated spatial artifacts and temporal flickering of generated videos. To achieve this, basing on a pretrained video diffusion model, we train a video ControlNet and inject it to the diffusion model as a condition on low frame-rate and low-resolution videos. To effectively train this video ControlNet, we design space-time data augmentation as well as video-aware conditioning. Benefiting from the above designs, VEnhancer yields to be stable during training and shares an elegant end-to-end training manner. Extensive experiments show that VEnhancer surpasses existing state-of-the-art video super-resolution and space-time super-resolution methods in enhancing AI-generated videos. Moreover, with VEnhancer, exisiting open-source state-of-the-art text-to-video method, VideoCrafter-2, reaches the top one in video generation benchmark -- VBench.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VEnhancer: Generative Space-Time Enhancement for Video Generation | TensorX