TensorX
返回文献探索

Paper · arXiv 2411.10510

SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers

Joseph Liu, Joshua Geddes, Ziyu Guo, Haomiao Jiang, Mahesh Kumar Nandwana

9 upvotesNovember 15, 2024arXiv 预印本
AI 摘要

SmoothCache accelerates DiT inference by caching key features, achieving significant speed-ups while maintaining quality across image, video, and audio tasks.

Diffusion TransformersDiTSmoothCacheinference accelerationattention modulesfeed-forward moduleslayer-wise representation errorsreal-time applications

Abstract

Diffusion Transformers (DiT) have emerged as powerful generative models for various tasks, including image, video, and speech synthesis. However, their inference process remains computationally expensive due to the repeated evaluation of resource-intensive attention and feed-forward modules. To address this, we introduce SmoothCache, a model-agnostic inference acceleration technique for DiT architectures. SmoothCache leverages the observed high similarity between layer outputs across adjacent diffusion timesteps. By analyzing layer-wise representation errors from a small calibration set, SmoothCache adaptively caches and reuses key features during inference. Our experiments demonstrate that SmoothCache achieves 8% to 71% speed up while maintaining or even improving generation quality across diverse modalities. We showcase its effectiveness on DiT-XL for image generation, Open-Sora for text-to-video, and Stable Audio Open for text-to-audio, highlighting its potential to enable real-time applications and broaden the accessibility of powerful DiT models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers | TensorX