TensorX
返回文献探索

Paper · arXiv 2310.01407

Conditional Diffusion Distillation

Kangfu Mei, Mauricio Delbracio, Hossein Talebi, Zhengzhong Tu, Vishal M. Patel, Peyman Milanfar

19 upvotesOctober 2, 2023arXiv 预印本
AI 摘要

A novel single-stage distillation method for generative diffusion models reduces sampling time while maintaining performance across tasks like super-resolution and image editing.

generative diffusion modelstext-to-image generationconditional generation tasksimage editingrestorationsuper-resolutiondiffusion priorsconditional samplingparameter-efficient distillationunconditional pre-trainingjoint-learningfine-tuned conditional diffusion models

Abstract

Generative diffusion models provide strong priors for text-to-image generation and thereby serve as a foundation for conditional generation tasks such as image editing, restoration, and super-resolution. However, one major limitation of diffusion models is their slow sampling time. To address this challenge, we present a novel conditional distillation method designed to supplement the diffusion priors with the help of image conditions, allowing for conditional sampling with very few steps. We directly distill the unconditional pre-training in a single stage through joint-learning, largely simplifying the previous two-stage procedures that involve both distillation and conditional finetuning separately. Furthermore, our method enables a new parameter-efficient distillation mechanism that distills each task with only a small number of additional parameters combined with the shared frozen unconditional backbone. Experiments across multiple tasks including super-resolution, image editing, and depth-to-image generation demonstrate that our method outperforms existing distillation techniques for the same sampling time. Notably, our method is the first distillation strategy that can match the performance of the much slower fine-tuned conditional diffusion models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号