TensorX
返回文献探索

Paper · arXiv 2311.10709

Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, Ishan Misra

25 upvotesNovember 17, 2023arXiv 预印本
AI 摘要

A text-to-video generation model that factors video creation into two steps—image generation followed by video generation—achieves superior quality and resolution without deep model cascades.

text-to-video generationdiffusionnoise schedulesmulti-stage trainingimage generationvideo generationhuman evaluationsImagen VideoPYOCOMake-A-VideoGen2Pika Labs

Abstract

We present Emu Video, a text-to-video generation model that factorizes the generation into two steps: first generating an image conditioned on the text, and then generating a video conditioned on the text and the generated image. We identify critical design decisions--adjusted noise schedules for diffusion, and multi-stage training--that enable us to directly generate high quality and high resolution videos, without requiring a deep cascade of models as in prior work. In human evaluations, our generated videos are strongly preferred in quality compared to all prior work--81% vs. Google's Imagen Video, 90% vs. Nvidia's PYOCO, and 96% vs. Meta's Make-A-Video. Our model outperforms commercial solutions such as RunwayML's Gen2 and Pika Labs. Finally, our factorizing approach naturally lends itself to animating images based on a user's text prompt, where our generations are preferred 96% over prior work.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号