TensorX
返回文献探索

Paper · arXiv 2404.11565

MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation

Kuan-Chieh, Wang, Daniil Ostashev, Yuwei Fang, Sergey Tulyakov, Kfir Aberman

15 upvotesApril 17, 2024arXiv 预印本
AI 摘要

Mixture-of-Attention (MoA) architecture personalizes text-to-image diffusion models by blending a fixed attention prior branch with a learnable personalized branch to enhance subject-context control.

Mixture-of-Attention (MoA)Mixture-of-Expertsdiffusion modelsattention pathwayspersonalized branchnon-personalized prior branchrouting mechanismhigh-quality imagessubject-context control

Abstract

We introduce a new architecture for personalization of text-to-image diffusion models, coined Mixture-of-Attention (MoA). Inspired by the Mixture-of-Experts mechanism utilized in large language models (LLMs), MoA distributes the generation workload between two attention pathways: a personalized branch and a non-personalized prior branch. MoA is designed to retain the original model's prior by fixing its attention layers in the prior branch, while minimally intervening in the generation process with the personalized branch that learns to embed subjects in the layout and context generated by the prior branch. A novel routing mechanism manages the distribution of pixels in each layer across these branches to optimize the blend of personalized and generic content creation. Once trained, MoA facilitates the creation of high-quality, personalized images featuring multiple subjects with compositions and interactions as diverse as those generated by the original model. Crucially, MoA enhances the distinction between the model's pre-existing capability and the newly augmented personalized intervention, thereby offering a more disentangled subject-context control that was previously unattainable. Project page: https://snap-research.github.io/mixture-of-attention

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation | TensorX