TensorX
返回文献探索

Paper · arXiv 2402.03162

Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion

Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xiaodong Chen, Jing Liao

20 upvotesFebruary 5, 2024arXiv 预印本
AI 摘要

Direct-a-Video enables independent control of object motion and camera movement in text-to-video generation using spatial and temporal cross-attention modulation.

text-to-video diffusion modelsobject motioncamera movementdecoupled controlspatial cross-attention modulationtemporal cross-attention layersself-supervised training

Abstract

Recent text-to-video diffusion models have achieved impressive progress. In practice, users often desire the ability to control object motion and camera movement independently for customized video creation. However, current methods lack the focus on separately controlling object motion and camera movement in a decoupled manner, which limits the controllability and flexibility of text-to-video models. In this paper, we introduce Direct-a-Video, a system that allows users to independently specify motions for one or multiple objects and/or camera movements, as if directing a video. We propose a simple yet effective strategy for the decoupled control of object motion and camera movement. Object motion is controlled through spatial cross-attention modulation using the model's inherent priors, requiring no additional optimization. For camera movement, we introduce new temporal cross-attention layers to interpret quantitative camera movement parameters. We further employ an augmentation-based approach to train these layers in a self-supervised manner on a small-scale dataset, eliminating the need for explicit motion annotation. Both components operate independently, allowing individual or combined control, and can generalize to open-domain scenarios. Extensive experiments demonstrate the superiority and effectiveness of our method. Project page: https://direct-a-video.github.io/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion | TensorX