TensorX
返回文献探索

Paper · arXiv 2412.14168

FashionComposer: Compositional Fashion Image Generation

Sihui Ji, Yiyang Wang, Xi Chen, Xiaogang Xu, Hao Luo, Hengshuang Zhao

17 upvotesDecember 18, 2024arXiv 预印本
AI 摘要

FashionComposer generates personalized fashion images using multi-modal input, featuring a universal framework, asset library, reference UNet, and subject-binding attention.

multi-modal inputtext promptparametric human modelgarment imageface imageuniversal frameworkscaled training dataasset libraryreference UNetsubject-binding attentionsemantic featureshuman album generationvirtual try-on tasks

Abstract

We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, and face image) and supports personalizing the appearance, pose, and figure of the human and assigning multiple garments in one pass. To achieve this, we first develop a universal framework capable of handling diverse input modalities. We construct scaled training data to enhance the model's robust compositional capabilities. To accommodate multiple reference images (garments and faces) seamlessly, we organize these references in a single image as an "asset library" and employ a reference UNet to extract appearance features. To inject the appearance features into the correct pixels in the generated result, we propose subject-binding attention. It binds the appearance features from different "assets" with the corresponding text features. In this way, the model could understand each asset according to their semantics, supporting arbitrary numbers and types of reference images. As a comprehensive solution, FashionComposer also supports many other applications like human album generation, diverse virtual try-on tasks, etc.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号