TensorX
返回文献探索

Paper · arXiv 2407.12705

IMAGDressing-v1: Customizable Virtual Dressing

Fei Shen, Xin Jiang, Xin He, Hu Ye, Cong Wang, Xiaoyu Du, Zechao Li, Jinghui Tang

13 upvotesJuly 17, 2024arXiv 预印本
AI 摘要

IMAGDressing-v1 enhances virtual try-on by generating editable human images with fixed garments using a garment UNet, hybrid attention module, and CLIP and VAE features, achieving state-of-the-art performance.

latent diffusion modelsVTONlocalized garment inpaintingvirtual dressing (VD) taskaffinity metric index (CAMI)garment UNetsemantic featurestexture featuresVAEhybrid attention modulefrozen self-attentiontrainable cross-attentiondenoising UNetControlNetIP-Adapterinteractive garment pairing (IGPair) datasethuman image synthesis

Abstract

Latest advances have achieved realistic virtual try-on (VTON) through localized garment inpainting using latent diffusion models, significantly enhancing consumers' online shopping experience. However, existing VTON technologies neglect the need for merchants to showcase garments comprehensively, including flexible control over garments, optional faces, poses, and scenes. To address this issue, we define a virtual dressing (VD) task focused on generating freely editable human images with fixed garments and optional conditions. Meanwhile, we design a comprehensive affinity metric index (CAMI) to evaluate the consistency between generated images and reference garments. Then, we propose IMAGDressing-v1, which incorporates a garment UNet that captures semantic features from CLIP and texture features from VAE. We present a hybrid attention module, including a frozen self-attention and a trainable cross-attention, to integrate garment features from the garment UNet into a frozen denoising UNet, ensuring users can control different scenes through text. IMAGDressing-v1 can be combined with other extension plugins, such as ControlNet and IP-Adapter, to enhance the diversity and controllability of generated images. Furthermore, to address the lack of data, we release the interactive garment pairing (IGPair) dataset, containing over 300,000 pairs of clothing and dressed images, and establish a standard pipeline for data assembly. Extensive experiments demonstrate that our IMAGDressing-v1 achieves state-of-the-art human image synthesis performance under various controlled conditions. The code and model will be available at https://github.com/muzishen/IMAGDressing.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号