TensorX
返回文献探索

Paper · arXiv 2401.11605

Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers

Katherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham, Daniel Z. Kaplan, Enrico Shippole

23 upvotesJanuary 21, 2024arXiv 预印本
AI 摘要

The Hourglass Diffusion Transformer achieves state-of-the-art performance on high-resolution image generation by combining the scalability of Transformers with the efficiency of U-Nets, operating directly in pixel space without additional high-resolution training techniques.

Hourglass Diffusion TransformerHDiTimage generative modelpixel-spaceTransformer architectureconvolutional U-Netsmultiscale architectureslatent autoencodersself-conditioningImageNetFFHQ

Abstract

We present the Hourglass Diffusion Transformer (HDiT), an image generative model that exhibits linear scaling with pixel count, supporting training at high-resolution (e.g. 1024 times 1024) directly in pixel-space. Building on the Transformer architecture, which is known to scale to billions of parameters, it bridges the gap between the efficiency of convolutional U-Nets and the scalability of Transformers. HDiT trains successfully without typical high-resolution training techniques such as multiscale architectures, latent autoencoders or self-conditioning. We demonstrate that HDiT performs competitively with existing models on ImageNet 256^2, and sets a new state-of-the-art for diffusion models on FFHQ-1024^2.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号