TensorX
返回文献探索

Paper · arXiv 2403.16627

SDXS: Real-Time One-Step Latent Diffusion Models with Image Conditions

Yuda Song, Zehao Sun, Xuanwu Yin

22 upvotesMarch 25, 2024arXiv 预印本
AI 摘要

A dual approach using model miniaturization and reduced sampling steps improves diffusion models' inference speed through knowledge distillation and feature matching.

diffusion modelsU-Netimage decoderknowledge distillationfeature matchingscore distillationSDXS-512SDXS-1024

Abstract

Recent advancements in diffusion models have positioned them at the forefront of image generation. Despite their superior performance, diffusion models are not without drawbacks; they are characterized by complex architectures and substantial computational demands, resulting in significant latency due to their iterative sampling process. To mitigate these limitations, we introduce a dual approach involving model miniaturization and a reduction in sampling steps, aimed at significantly decreasing model latency. Our methodology leverages knowledge distillation to streamline the U-Net and image decoder architectures, and introduces an innovative one-step DM training technique that utilizes feature matching and score distillation. We present two models, SDXS-512 and SDXS-1024, achieving inference speeds of approximately 100 FPS (30x faster than SD v1.5) and 30 FP (60x faster than SDXL) on a single GPU, respectively. Moreover, our training approach offers promising applications in image-conditioned control, facilitating efficient image-to-image translation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SDXS: Real-Time One-Step Latent Diffusion Models with Image Conditions | TensorX