TensorX
返回文献探索

Paper · arXiv 2312.00845

VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Models

Hyeonho Jeong, Geon Yeong Park, Jong Chul Ye

39 upvotesDecember 1, 2023arXiv 预印本
AI 摘要

The VMC framework enhances video generation by tuning temporal attention layers in diffusion models to accurately reproduce and diversify motion with a motion distillation objective.

text-to-video diffusion modelsvideo diffusion modelstemporal attention layersmotion customizationmotion distillationresidual vectorslow-frequency motion trajectorieshigh-frequency motion-unrelated noisestate-of-the-art video generative models

Abstract

Text-to-video diffusion models have advanced video generation significantly. However, customizing these models to generate videos with tailored motions presents a substantial challenge. In specific, they encounter hurdles in (a) accurately reproducing motion from a target video, and (b) creating diverse visual variations. For example, straightforward extensions of static image customization methods to video often lead to intricate entanglements of appearance and motion data. To tackle this, here we present the Video Motion Customization (VMC) framework, a novel one-shot tuning approach crafted to adapt temporal attention layers within video diffusion models. Our approach introduces a novel motion distillation objective using residual vectors between consecutive frames as a motion reference. The diffusion process then preserves low-frequency motion trajectories while mitigating high-frequency motion-unrelated noise in image space. We validate our method against state-of-the-art video generative models across diverse real-world motions and contexts. Our codes, data and the project demo can be found at https://video-motion-customization.github.io

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Models | TensorX