TensorX
返回文献探索

Paper · arXiv 2401.12945

Lumiere: A Space-Time Diffusion Model for Video Generation

Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Herrmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Yuanzhen Li, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, Inbar Mosseri

86 upvotesJanuary 23, 2024arXiv 预印本
AI 摘要

A text-to-video diffusion model using Space-Time U-Net architecture generates realistic, diverse, and coherent videos through a single pass, achieving state-of-the-art results and supporting various content creation and editing tasks.

Space-Time U-Netdiffusion modeltemporal durationspatial down- and up-samplingtemporal down- and up-samplingtext-to-image diffusion modelfull-frame-ratelow-resolution videoimage-to-videovideo inpaintingstylized generation

Abstract

We introduce Lumiere -- a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion -- a pivotal challenge in video synthesis. To this end, we introduce a Space-Time U-Net architecture that generates the entire temporal duration of the video at once, through a single pass in the model. This is in contrast to existing video models which synthesize distant keyframes followed by temporal super-resolution -- an approach that inherently makes global temporal consistency difficult to achieve. By deploying both spatial and (importantly) temporal down- and up-sampling and leveraging a pre-trained text-to-image diffusion model, our model learns to directly generate a full-frame-rate, low-resolution video by processing it in multiple space-time scales. We demonstrate state-of-the-art text-to-video generation results, and show that our design easily facilitates a wide range of content creation tasks and video editing applications, including image-to-video, video inpainting, and stylized generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Lumiere: A Space-Time Diffusion Model for Video Generation | TensorX