TensorX
返回文献探索

Paper · arXiv 2408.08189

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

Jiasong Feng, Ao Ma, Jing Wang, Bo Cheng, Xiaodan Liang, Dawei Leng, Yuhui Yin

18 upvotesAugust 15, 2024arXiv 预印本
AI 摘要

FancyVideo introduces the CTGM to improve temporal consistency in text-to-video synthesis by incorporating TII, TAR, and TFB for frame-specific text control.

Cross-frame Textual Guidance ModuleCTGMTemporal Information InjectorTIITemporal Affinity RefinerTARTemporal Feature BoosterTFBtext-to-videoT2Vlatency featurescross-attentionEvalCrafter benchmark

Abstract

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations without frame-specific textual guidance. Thus, the model's capacity to comprehend the temporal logic conveyed in prompts and generate videos with coherent motion is restricted. To tackle this limitation, we introduce FancyVideo, an innovative video generator that improves the existing text-control mechanism with the well-designed Cross-frame Textual Guidance Module (CTGM). Specifically, CTGM incorporates the Temporal Information Injector (TII), Temporal Affinity Refiner (TAR), and Temporal Feature Booster (TFB) at the beginning, middle, and end of cross-attention, respectively, to achieve frame-specific textual guidance. Firstly, TII injects frame-specific information from latent features into text conditions, thereby obtaining cross-frame textual conditions. Then, TAR refines the correlation matrix between cross-frame textual conditions and latent features along the time dimension. Lastly, TFB boosts the temporal consistency of latent features. Extensive experiments comprising both quantitative and qualitative evaluations demonstrate the effectiveness of FancyVideo. Our approach achieves state-of-the-art T2V generation results on the EvalCrafter benchmark and facilitates the synthesis of dynamic and consistent videos. The video show results can be available at https://fancyvideo.github.io/, and we will make our code and model weights publicly available.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance | TensorX