TensorX
返回文献探索

Paper · arXiv 2508.07901

Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation

Bowen Xue, Qixin Yan, Wenjing Wang, Hao Liu, Chen Li

40 upvotesAugust 11, 2025arXiv 预印本
AI 摘要

A lightweight framework for identity preservation in video generation using conditional image branches and restricted self-attentions outperforms full-parameter methods with minimal additional parameters.

conditional image branchpre-trained video generation modelrestricted self-attentionsconditional position mappingsubject-driven video generationpose-referenced video generationstylizationface swapping

Abstract

Generating high-fidelity human videos that match user-specified identities is important yet challenging in the field of generative AI. Existing methods often rely on an excessive number of training parameters and lack compatibility with other AIGC tools. In this paper, we propose Stand-In, a lightweight and plug-and-play framework for identity preservation in video generation. Specifically, we introduce a conditional image branch into the pre-trained video generation model. Identity control is achieved through restricted self-attentions with conditional position mapping, and can be learned quickly with only 2000 pairs. Despite incorporating and training just sim1\% additional parameters, our framework achieves excellent results in video quality and identity preservation, outperforming other full-parameter training methods. Moreover, our framework can be seamlessly integrated for other tasks, such as subject-driven video generation, pose-referenced video generation, stylization, and face swapping.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation | TensorX