TensorX
返回文献探索

Paper · arXiv 2409.17565

Pixel-Space Post-Training of Latent Diffusion Models

Christina Zhang, Simran Motwani, Matthew Yu, Ji Hou, Felix Juefei-Xu, Sam Tsai, Peter Vajda, Zijian He, Jialiang Wang

20 upvotesSeptember 26, 2024arXiv 预印本
AI 摘要

Adding pixel-space supervision to latent diffusion models improves high-frequency detail preservation during post-training without compromising text alignment quality.

latent diffusion modelspixel-space supervisionhigh-frequency detailslatent spaceDiT transformerU-Net diffusion modelsvisual qualityvisual flaw metricstext alignment quality

Abstract

Latent diffusion models (LDMs) have made significant advancements in the field of image generation in recent years. One major advantage of LDMs is their ability to operate in a compressed latent space, allowing for more efficient training and deployment. However, despite these advantages, challenges with LDMs still remain. For example, it has been observed that LDMs often generate high-frequency details and complex compositions imperfectly. We hypothesize that one reason for these flaws is due to the fact that all pre- and post-training of LDMs are done in latent space, which is typically 8 times 8 lower spatial-resolution than the output images. To address this issue, we propose adding pixel-space supervision in the post-training process to better preserve high-frequency details. Experimentally, we show that adding a pixel-space objective significantly improves both supervised quality fine-tuning and preference-based post-training by a large margin on a state-of-the-art DiT transformer and U-Net diffusion models in both visual quality and visual flaw metrics, while maintaining the same text alignment quality.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Pixel-Space Post-Training of Latent Diffusion Models | TensorX