TensorX
返回文献探索

Paper · arXiv 2401.08740

SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers

Nanye Ma, Mark Goldstein, Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden, Saining Xie

13 upvotesJanuary 16, 2024arXiv 预印本
AI 摘要

Scalable Interpolant Transformers (SiT), built on Diffusion Transformers (DiT), enhance generative models through a flexible interpolant framework, achieving superior performance on ImageNet 256x256 with tunable diffusion coefficients.

Scalable Interpolant TransformersSiTDiffusion TransformersDiTinterpolant frameworkdynamical transportdiscrete time learningcontinuous time learningstochastic samplerdiffusion coefficientsFID-50K score

Abstract

We present Scalable Interpolant Transformers (SiT), a family of generative models built on the backbone of Diffusion Transformers (DiT). The interpolant framework, which allows for connecting two distributions in a more flexible way than standard diffusion models, makes possible a modular study of various design choices impacting generative models built on dynamical transport: using discrete vs. continuous time learning, deciding the objective for the model to learn, choosing the interpolant connecting the distributions, and deploying a deterministic or stochastic sampler. By carefully introducing the above ingredients, SiT surpasses DiT uniformly across model sizes on the conditional ImageNet 256x256 benchmark using the exact same backbone, number of parameters, and GFLOPs. By exploring various diffusion coefficients, which can be tuned separately from learning, SiT achieves an FID-50K score of 2.06.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers | TensorX