TensorX
返回文献探索

Paper · arXiv 2503.16397

Scale-wise Distillation of Diffusion Models

Nikita Starodubcev, Denis Kuznedelev, Artem Babenko, Dmitry Baranchuk

42 upvotesMarch 20, 2025arXiv 预印本
AI 摘要

A scale-wise distillation framework for diffusion models reduces computational costs and improves inference times by incorporating next-scale predictions and enhancing distribution matching methods.

diffusion modelsscale-wise distillationnext-scale predictionimplicit spectral autoregressiondenoisingcomputational costsdistribution matchingpatch lossinference timestext-to-image diffusion models

Abstract

We present SwD, a scale-wise distillation framework for diffusion models (DMs), which effectively employs next-scale prediction ideas for diffusion-based few-step generators. In more detail, SwD is inspired by the recent insights relating diffusion processes to the implicit spectral autoregression. We suppose that DMs can initiate generation at lower data resolutions and gradually upscale the samples at each denoising step without loss in performance while significantly reducing computational costs. SwD naturally integrates this idea into existing diffusion distillation methods based on distribution matching. Also, we enrich the family of distribution matching approaches by introducing a novel patch loss enforcing finer-grained similarity to the target distribution. When applied to state-of-the-art text-to-image diffusion models, SwD approaches the inference times of two full resolution steps and significantly outperforms the counterparts under the same computation budget, as evidenced by automated metrics and human preference studies.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Scale-wise Distillation of Diffusion Models | TensorX