TensorX
返回文献探索

Paper · arXiv 2411.10836

AnimateAnything: Consistent and Controllable Animation for Video Generation

Guojun Lei, Chi Wang, Hong Li, Rong Zhang, Yikai Wang, Weiwei Xu

24 upvotesNovember 16, 2024arXiv 预印本
AI 摘要

A unified controllable video generation method AnimateAnything utilizes a multi-scale fusion network and optical flows to handle various conditions, achieving high consistency and stability in generated videos.

multi-scale control feature fusion networkframe-by-frame optical flowsmotion priorsfrequency-based stabilization moduletemporal coherencefrequency domain consistency

Abstract

We present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions, including camera trajectories, text prompts, and user motion annotations. Specifically, we carefully design a multi-scale control feature fusion network to construct a common motion representation for different conditions. It explicitly converts all control information into frame-by-frame optical flows. Then we incorporate the optical flows as motion priors to guide final video generation. In addition, to reduce the flickering issues caused by large-scale motion, we propose a frequency-based stabilization module. It can enhance temporal coherence by ensuring the video's frequency domain consistency. Experiments demonstrate that our method outperforms the state-of-the-art approaches. For more details and videos, please refer to the webpage: https://yu-shaonian.github.io/Animate_Anything/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号