TensorX
返回文献探索

Paper · arXiv 2406.17758

MotionBooth: Motion-Aware Customized Text-to-Video Generation

Jianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang, Qianyu Zhou, Yining Li, Yunhai Tong, Kai Chen

19 upvotesJune 25, 2024arXiv 预印本
AI 摘要

MotionBooth fine-tunes a text-to-video model using a few images to animate customized subjects with precise control over object and camera movements, enhancing performance with subject region loss, video preservation loss, and subject token cross-attention loss.

text-to-video modelsubject region lossvideo preservation losssubject token cross-attention losscross-attention map manipulationlatent shift module

Abstract

In this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we efficiently fine-tune a text-to-video model to capture the object's shape and attributes accurately. Our approach presents subject region loss and video preservation loss to enhance the subject's learning performance, along with a subject token cross-attention loss to integrate the customized subject with motion control signals. Additionally, we propose training-free techniques for managing subject and camera motions during inference. In particular, we utilize cross-attention map manipulation to govern subject motion and introduce a novel latent shift module for camera movement control as well. MotionBooth excels in preserving the appearance of subjects while simultaneously controlling the motions in generated videos. Extensive quantitative and qualitative evaluations demonstrate the superiority and effectiveness of our method. Our project page is at https://jianzongwu.github.io/projects/motionbooth

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MotionBooth: Motion-Aware Customized Text-to-Video Generation | TensorX