TensorX
返回文献探索

Paper · arXiv 2411.02395

Training-free Regional Prompting for Diffusion Transformers

Anthony Chen, Jianjin Xu, Wenzhao Zheng, Gaole Dai, Yida Wang, Renrui Zhang, Haofan Wang, Shanghang Zhang

25 upvotesNovember 4, 2024arXiv 预印本
AI 摘要

Regional prompting is implemented for the FLUX diffusion model using attention manipulation to achieve fine-grained text-to-image generation without training.

diffusion modelstext-to-image generationlarge language modelsT5Llamaregional promptingUNetDiffusion TransformerDiTattention manipulationfined-grained compositional text-to-image generation

Abstract

Diffusion models have demonstrated excellent capabilities in text-to-image generation. Their semantic understanding (i.e., prompt following) ability has also been greatly improved with large language models (e.g., T5, Llama). However, existing models cannot perfectly handle long and complex text prompts, especially when the text prompts contain various objects with numerous attributes and interrelated spatial relationships. While many regional prompting methods have been proposed for UNet-based models (SD1.5, SDXL), but there are still no implementations based on the recent Diffusion Transformer (DiT) architecture, such as SD3 and FLUX.1.In this report, we propose and implement regional prompting for FLUX.1 based on attention manipulation, which enables DiT with fined-grained compositional text-to-image generation capability in a training-free manner. Code is available at https://github.com/antonioo-c/Regional-Prompting-FLUX.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Training-free Regional Prompting for Diffusion Transformers | TensorX