TensorX
返回文献探索

Paper · arXiv 2312.06971

CCM: Adding Conditional Controls to Text-to-Image Consistency Models

Jie Xiao, Kai Zhu, Han Zhang, Zhiheng Liu, Yujun Shen, Yu Liu, Xueyang Fu, Zheng-Jun Zha

11 upvotesDecember 12, 2023arXiv 预印本
AI 摘要

Research explores methods to integrate ControlNet-like conditional control into Consistency Models, finding approaches for high-level semantic control, alternative training methods, and lightweight adapters.

Consistency ModelsControlNetdiffusion modelshigh-level semantic controlslow-level detail controlConsistency Traininglightweight adapteredgedepthhuman poselow-resolution imagemasked imagetext-to-image latent consistency models

Abstract

Consistency Models (CMs) have showed a promise in creating visual content efficiently and with high quality. However, the way to add new conditional controls to the pretrained CMs has not been explored. In this technical report, we consider alternative strategies for adding ControlNet-like conditional control to CMs and present three significant findings. 1) ControlNet trained for diffusion models (DMs) can be directly applied to CMs for high-level semantic controls but struggles with low-level detail and realism control. 2) CMs serve as an independent class of generative models, based on which ControlNet can be trained from scratch using Consistency Training proposed by Song et al. 3) A lightweight adapter can be jointly optimized under multiple conditions through Consistency Training, allowing for the swift transfer of DMs-based ControlNet to CMs. We study these three solutions across various conditional controls, including edge, depth, human pose, low-resolution image and masked image with text-to-image latent consistency models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
CCM: Adding Conditional Controls to Text-to-Image Consistency Models | TensorX