TensorX
返回文献探索

Paper · arXiv 2501.03931

Magic Mirror: ID-Preserved Video Generation in Video Diffusion Transformers

Yuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng, Zexin Yan, Eric Lo, Jiaya Jia

15 upvotesJanuary 7, 2025arXiv 预印本
AI 摘要

Magic Mirror framework uses Video Diffusion Transformers to generate identity-preserved videos with natural motion and high quality, overcoming limitations of existing methods.

Video Diffusion Transformersdual-branch facial feature extractoridentity and structural featureslightweight cross-modal adapterConditioned Adaptive Normalizationtwo-stage training strategysynthetic identity pairs

Abstract

We present Magic Mirror, a framework for generating identity-preserved videos with cinematic-level quality and dynamic motion. While recent advances in video diffusion models have shown impressive capabilities in text-to-video generation, maintaining consistent identity while producing natural motion remains challenging. Previous methods either require person-specific fine-tuning or struggle to balance identity preservation with motion diversity. Built upon Video Diffusion Transformers, our method introduces three key components: (1) a dual-branch facial feature extractor that captures both identity and structural features, (2) a lightweight cross-modal adapter with Conditioned Adaptive Normalization for efficient identity integration, and (3) a two-stage training strategy combining synthetic identity pairs with video data. Extensive experiments demonstrate that Magic Mirror effectively balances identity consistency with natural motion, outperforming existing methods across multiple metrics while requiring minimal parameters added. The code and model will be made publicly available at: https://github.com/dvlab-research/MagicMirror/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号