TensorX
返回文献探索

Paper · arXiv 2401.04468

MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation

Weimin Wang, Jiawei Liu, Zhijie Lin, Jiangqiao Yan, Shuo Chen, Chetwin Low, Tuyen Hoang, Jie Wu, Jun Hao Liew, Hanshu Yan, Daquan Zhou, Jiashi Feng

49 upvotesJanuary 9, 2024arXiv 预印本
AI 摘要

MagicVideo-V2 generates high-fidelity and smooth videos from text using an integrated pipeline that includes text-to-image, video motion generation, and frame interpolation modules, outperforming existing systems in user evaluations.

text-to-imagevideo motion generatorreference image embeddingframe interpolationvideo generation pipelineuser evaluation

Abstract

The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that integrates the text-to-image model, video motion generator, reference image embedding module and frame interpolation module into an end-to-end video generation pipeline. Benefiting from these architecture designs, MagicVideo-V2 can generate an aesthetically pleasing, high-resolution video with remarkable fidelity and smoothness. It demonstrates superior performance over leading Text-to-Video systems such as Runway, Pika 1.0, Morph, Moon Valley and Stable Video Diffusion model via user evaluation at large scale.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation | TensorX