TensorX
返回文献探索

Paper · arXiv 2312.06662

Photorealistic Video Generation with Diffusion Models

Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, José Lezama

24 upvotesDecember 11, 2023arXiv 预印本
AI 摘要

A transformer-based diffusion model using causal encoding and window attention generates high-resolution, photorealistic videos, achieving state-of-the-art performance without classifier-free guidance and includes a cascade for text-to-video generation.

transformer-based approachdiffusion modelingcausal encoderunified latent spacejoint spatialspatiotemporal generative modelingwindow attention architecturevideo generationImageNetUCF-101Kinetics-600latent video diffusion modelvideo super-resolution diffusion modelstext-to-video generation

Abstract

We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encoder to jointly compress images and videos within a unified latent space, enabling training and generation across modalities. Second, for memory and training efficiency, we use a window attention architecture tailored for joint spatial and spatiotemporal generative modeling. Taken together these design decisions enable us to achieve state-of-the-art performance on established video (UCF-101 and Kinetics-600) and image (ImageNet) generation benchmarks without using classifier free guidance. Finally, we also train a cascade of three models for the task of text-to-video generation consisting of a base latent video diffusion model, and two video super-resolution diffusion models to generate videos of 512 times 896 resolution at 8 frames per second.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Photorealistic Video Generation with Diffusion Models | TensorX