TensorX
返回文献探索

Paper · arXiv 2403.01800

AtomoVideo: High Fidelity Image-to-Video Generation

Litong Gong, Yiran Zhu, Weijie Li, Xiaoyang Kang, Biao Wang, Tiezheng Ge, Bo Zheng

22 upvotesMarch 4, 2024arXiv 预印本
AI 摘要

AtomoVideo generates high-fidelity videos from images using multi-granularity injection, achieving superior motion intensity and temporal stability.

multi-granularity image injectiontemporal consistencyframe predictioniterative generationadapter trainingpersonalized modelscontrollable modules

Abstract

Recently, video generation has achieved significant rapid development based on superior text-to-image generation techniques. In this work, we propose a high fidelity framework for image-to-video generation, named AtomoVideo. Based on multi-granularity image injection, we achieve higher fidelity of the generated video to the given image. In addition, thanks to high quality datasets and training strategies, we achieve greater motion intensity while maintaining superior temporal consistency and stability. Our architecture extends flexibly to the video frame prediction task, enabling long sequence prediction through iterative generation. Furthermore, due to the design of adapter training, our approach can be well combined with existing personalised models and controllable modules. By quantitatively and qualitatively evaluation, AtomoVideo achieves superior results compared to popular methods, more examples can be found on our project website: https://atomo- video.github.io/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
AtomoVideo: High Fidelity Image-to-Video Generation | TensorX