TensorX
返回文献探索

Paper · arXiv 2306.05427

Grounded Text-to-Image Synthesis with Attention Refocusing

Quynh Phung, Songwei Ge, Jia-Bin Huang

3 upvotesJune 8, 2023arXiv 预印本
AI 摘要

Two novel losses improve attention focus in diffusion models, enhancing text-to-image alignment on benchmarks.

diffusion modelstext-to-image synthesiscross-attentionself-attentionattention mapslayoutDrawBenchHRS benchmarksLarge Language Models

Abstract

Driven by scalable diffusion models trained on large-scale paired text-image datasets, text-to-image synthesis methods have shown compelling results. However, these models still fail to precisely follow the text prompt when multiple objects, attributes, and spatial compositions are involved in the prompt. In this paper, we identify the potential reasons in both the cross-attention and self-attention layers of the diffusion model. We propose two novel losses to refocus the attention maps according to a given layout during the sampling process. We perform comprehensive experiments on the DrawBench and HRS benchmarks using layouts synthesized by Large Language Models, showing that our proposed losses can be integrated easily and effectively into existing text-to-image methods and consistently improve their alignment between the generated images and the text prompts.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Grounded Text-to-Image Synthesis with Attention Refocusing | TensorX