TensorX
返回文献探索

Paper · arXiv 2307.01952

SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, Robin Rombach

93 upvotesJuly 4, 2023arXiv 预印本
AI 摘要

SDXL, a latent diffusion model using a larger UNet with additional text encoders and attention mechanisms, improves text-to-image synthesis significantly.

latent diffusion modelUNetattention blockscross-attention contexttext encoderconditioning schemesrefinement modelimage-to-image techniquevisual fidelity

Abstract

We present SDXL, a latent diffusion model for text-to-image synthesis. Compared to previous versions of Stable Diffusion, SDXL leverages a three times larger UNet backbone: The increase of model parameters is mainly due to more attention blocks and a larger cross-attention context as SDXL uses a second text encoder. We design multiple novel conditioning schemes and train SDXL on multiple aspect ratios. We also introduce a refinement model which is used to improve the visual fidelity of samples generated by SDXL using a post-hoc image-to-image technique. We demonstrate that SDXL shows drastically improved performance compared the previous versions of Stable Diffusion and achieves results competitive with those of black-box state-of-the-art image generators. In the spirit of promoting open research and fostering transparency in large model training and evaluation, we provide access to code and model weights at https://github.com/Stability-AI/generative-models

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis | TensorX