TensorX
返回文献探索

Paper · arXiv 2410.08159

DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation

Jiatao Gu, Yuyang Wang, Yizhe Zhang, Qihang Zhang, Dinghuai Zhang, Navdeep Jaitly, Josh Susskind, Shuangfei Zhai

26 upvotesOctober 10, 2024arXiv 预印本
AI 摘要

DART, a transformer-based model combining autoregressive and diffusion components in a non-Markovian framework, offers competitive performance on image and text-to-image generation tasks.

diffusion modelsMarkovian processautoregressive (AR) modelsnon-Markovian frameworkimage patchesspectral denoisingimage quantizationclass-conditioned generationtext-to-image generation

Abstract

Diffusion models have become the dominant approach for visual generation. They are trained by denoising a Markovian process that gradually adds noise to the input. We argue that the Markovian property limits the models ability to fully utilize the generation trajectory, leading to inefficiencies during training and inference. In this paper, we propose DART, a transformer-based model that unifies autoregressive (AR) and diffusion within a non-Markovian framework. DART iteratively denoises image patches spatially and spectrally using an AR model with the same architecture as standard language models. DART does not rely on image quantization, enabling more effective image modeling while maintaining flexibility. Furthermore, DART seamlessly trains with both text and image data in a unified model. Our approach demonstrates competitive performance on class-conditioned and text-to-image generation tasks, offering a scalable, efficient alternative to traditional diffusion models. Through this unified framework, DART sets a new benchmark for scalable, high-quality image synthesis.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation | TensorX