TensorX
返回文献探索

Paper · arXiv 2306.00980

SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds

Yanyu Li, Huan Wang, Qing Jin, Ju Hu, Pavlo Chemerys, Yun Fu, Yanzhi Wang, Sergey Tulyakov, Jian Ren

16 upvotesJune 1, 2023arXiv 预印本
AI 摘要

Text-to-image diffusion models are optimized for mobile devices through efficient architecture and improved distillation techniques, achieving high-quality image generation with reduced computational requirements.

diffusion modelstext-to-imageUNetdenoising stepsFIDCLIP scoresclassifier-free guidancedata distillationstep distillationMS-COCO

Abstract

Text-to-image diffusion models can create stunning images from natural language descriptions that rival the work of professional artists and photographers. However, these models are large, with complex network architectures and tens of denoising iterations, making them computationally expensive and slow to run. As a result, high-end GPUs and cloud-based inference are required to run diffusion models at scale. This is costly and has privacy implications, especially when user data is sent to a third party. To overcome these challenges, we present a generic approach that, for the first time, unlocks running text-to-image diffusion models on mobile devices in less than 2 seconds. We achieve so by introducing efficient network architecture and improving step distillation. Specifically, we propose an efficient UNet by identifying the redundancy of the original model and reducing the computation of the image decoder via data distillation. Further, we enhance the step distillation by exploring training strategies and introducing regularization from classifier-free guidance. Our extensive experiments on MS-COCO show that our model with 8 denoising steps achieves better FID and CLIP scores than Stable Diffusion v1.5 with 50 steps. Our work democratizes content creation by bringing powerful text-to-image diffusion models to the hands of users.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds | TensorX