TensorX
返回文献探索

Paper · arXiv 2312.02139

DiffiT: Diffusion Vision Transformers for Image Generation

Ali Hatamizadeh, Jiaming Song, Guilin Liu, Jan Kautz, Arash Vahdat

15 upvotesDecember 4, 2023arXiv 预印本
AI 摘要

A novel diffusion model using vision transformers with a hierarchical architecture and time-dependent self-attention achieves state-of-the-art performance in image generation.

diffusion modelsdenoising neural networkconvolutional residual U-Netsvision transformersDiffusion Vision TransformersDiffiTU-shaped encoderdecodertime-dependent self-attentionlatent DiffiThigh-resolution image generationclass-conditional synthesisunconditional synthesisFID score

Abstract

Diffusion models with their powerful expressivity and high sample quality have enabled many new applications and use-cases in various domains. For sample generation, these models rely on a denoising neural network that generates images by iterative denoising. Yet, the role of denoising network architecture is not well-studied with most efforts relying on convolutional residual U-Nets. In this paper, we study the effectiveness of vision transformers in diffusion-based generative learning. Specifically, we propose a new model, denoted as Diffusion Vision Transformers (DiffiT), which consists of a hybrid hierarchical architecture with a U-shaped encoder and decoder. We introduce a novel time-dependent self-attention module that allows attention layers to adapt their behavior at different stages of the denoising process in an efficient manner. We also introduce latent DiffiT which consists of transformer model with the proposed self-attention layers, for high-resolution image generation. Our results show that DiffiT is surprisingly effective in generating high-fidelity images, and it achieves state-of-the-art (SOTA) benchmarks on a variety of class-conditional and unconditional synthesis tasks. In the latent space, DiffiT achieves a new SOTA FID score of 1.73 on ImageNet-256 dataset. Repository: https://github.com/NVlabs/DiffiT

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号