TensorX
返回文献探索

Paper · arXiv 2504.17816

Subject-driven Video Generation via Disentangled Identity and Motion

Daneul Kim, Jingxu Zhang, Wonjoon Jin, Sunghyun Cho, Qi Dai, Jaesik Park, Chong Luo

13 upvotesApril 23, 2025arXiv 预印本
AI 摘要

A method for zero-shot video customization achieves strong performance by decoupling subject-specific learning from temporal dynamics using image customization datasets and randomization techniques.

identity injectionimage customization datasettemporal modelingimage-to-video trainingrandom image token droppingstochastic switchingcatastrophic forgettingsubject consistencyscalabilityzero-shot settings

Abstract

We propose to train a subject-driven customized video generation model through decoupling the subject-specific learning from temporal dynamics in zero-shot without additional tuning. A traditional method for video customization that is tuning-free often relies on large, annotated video datasets, which are computationally expensive and require extensive annotation. In contrast to the previous approach, we introduce the use of an image customization dataset directly on training video customization models, factorizing the video customization into two folds: (1) identity injection through image customization dataset and (2) temporal modeling preservation with a small set of unannotated videos through the image-to-video training method. Additionally, we employ random image token dropping with randomized image initialization during image-to-video fine-tuning to mitigate the copy-and-paste issue. To further enhance learning, we introduce stochastic switching during joint optimization of subject-specific and temporal features, mitigating catastrophic forgetting. Our method achieves strong subject consistency and scalability, outperforming existing video customization models in zero-shot settings, demonstrating the effectiveness of our framework.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号