TensorX
返回文献探索

Paper · arXiv 2404.19427

InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation

Chanran Kim, Jeongin Lee, Shichang Joung, Bongmo Kim, Yeul-Min Baek

74 upvotesApril 30, 2024arXiv 预印本
AI 摘要

InstantFamily uses masked cross-attention and multimodal embeddings to generate images with preserved multi-IDs and high visual coherence.

masked cross-attentionmultimodal embedding stackpre-trained face recognition modelmulti-ID image generationsingle-ID preservationstate-of-the-art performancescalability

Abstract

In the field of personalized image generation, the ability to create images preserving concepts has significantly improved. Creating an image that naturally integrates multiple concepts in a cohesive and visually appealing composition can indeed be challenging. This paper introduces "InstantFamily," an approach that employs a novel masked cross-attention mechanism and a multimodal embedding stack to achieve zero-shot multi-ID image generation. Our method effectively preserves ID as it utilizes global and local features from a pre-trained face recognition model integrated with text conditions. Additionally, our masked cross-attention mechanism enables the precise control of multi-ID and composition in the generated images. We demonstrate the effectiveness of InstantFamily through experiments showing its dominance in generating images with multi-ID, while resolving well-known multi-ID generation problems. Additionally, our model achieves state-of-the-art performance in both single-ID and multi-ID preservation. Furthermore, our model exhibits remarkable scalability with a greater number of ID preservation than it was originally trained with.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation | TensorX