TensorX
返回文献探索

Paper · arXiv 2602.17270

Unified Latents (UL): How to train your latents

Jonathan Heek, Emiel Hoogeboom, Thomas Mensink, Tim Salimans

63 upvotesFebruary 19, 2026arXiv 预印本
AI 摘要

Unified Latents framework learns joint latent representations using diffusion prior regularization and diffusion model decoding, achieving competitive FID scores with reduced training compute.

diffusion priordiffusion modellatent representationstraining objectivelatent bitrateFIDPSNRFLOPsFVD

Abstract

We present Unified Latents (UL), a framework for learning latent representations that are jointly regularized by a diffusion prior and decoded by a diffusion model. By linking the encoder's output noise to the prior's minimum noise level, we obtain a simple training objective that provides a tight upper bound on the latent bitrate. On ImageNet-512, our approach achieves competitive FID of 1.4, with high reconstruction quality (PSNR) while requiring fewer training FLOPs than models trained on Stable Diffusion latents. On Kinetics-600, we set a new state-of-the-art FVD of 1.3.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Unified Latents (UL): How to train your latents | TensorX