TensorX
返回文献探索

Paper · arXiv 2410.11081

Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

Cheng Lu, Yang Song

18 upvotesOctober 14, 2024arXiv 预印本
AI 摘要

A new theoretical framework and algorithm improve continuous-time consistency models, enabling large-scale training and achieving near-state-of-the-art FID scores.

consistency modelsdiffusion-based generative modelsdiscretized timestepscontinuous-time formulationstraining instabilitydiffusion process parameterizationnetwork architecturetraining objectivesImageNetCIFAR-10FID scores

Abstract

Consistency models (CMs) are a powerful class of diffusion-based generative models optimized for fast sampling. Most existing CMs are trained using discretized timesteps, which introduce additional hyperparameters and are prone to discretization errors. While continuous-time formulations can mitigate these issues, their success has been limited by training instability. To address this, we propose a simplified theoretical framework that unifies previous parameterizations of diffusion models and CMs, identifying the root causes of instability. Based on this analysis, we introduce key improvements in diffusion process parameterization, network architecture, and training objectives. These changes enable us to train continuous-time CMs at an unprecedented scale, reaching 1.5B parameters on ImageNet 512x512. Our proposed training algorithm, using only two sampling steps, achieves FID scores of 2.06 on CIFAR-10, 1.48 on ImageNet 64x64, and 1.88 on ImageNet 512x512, narrowing the gap in FID scores with the best existing diffusion models to within 10%.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models | TensorX