TensorX
返回文献探索

Paper · arXiv 2412.01819

Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis

Anton Voronov, Denis Kuznedelev, Mikhail Khoroshikh, Valentin Khrulkov, Dmitry Baranchuk

34 upvotesDecember 2, 2024arXiv 预印本
AI 摘要

Switti, a scale-wise transformer for text-to-image generation, improves convergence and performance through architectural modifications, reduces memory usage, and achieves competitive results compared to diffusion models with significant speed advantages.

scale-wise transformertext-to-image generationnext-scale predictionAR modelsself-attention mapspretrained scale-wise AR modelnon-AR counterpartclassifier-free guidancehigh-resolution scalessampling accelerationhuman preference studiesautomated evaluationsdiffusion models

Abstract

This work presents Switti, a scale-wise transformer for text-to-image generation. Starting from existing next-scale prediction AR models, we first explore them for T2I generation and propose architectural modifications to improve their convergence and overall performance. We then observe that self-attention maps of our pretrained scale-wise AR model exhibit weak dependence on preceding scales. Based on this insight, we propose a non-AR counterpart facilitating {sim}11% faster sampling and lower memory usage while also achieving slightly better generation quality.Furthermore, we reveal that classifier-free guidance at high-resolution scales is often unnecessary and can even degrade performance. %may be not only unnecessary but potentially detrimental. By disabling guidance at these scales, we achieve an additional sampling acceleration of {sim}20% and improve the generation of fine-grained details. Extensive human preference studies and automated evaluations show that Switti outperforms existing T2I AR models and competes with state-of-the-art T2I diffusion models while being up to 7{times} faster.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号