TensorX
返回文献探索

Paper · arXiv 2605.27102

JLT: Clean-Latent Prediction in Latent Diffusion Transformers

Funing Fu, Tenghui Wang, Junyong Cen, Qichao Zhu, Guanyu Zhou

33 upvotesMay 26, 2026arXiv 预印本
AI 摘要

Latent diffusion models using clean-data prediction outperform velocity prediction in compressed representations, demonstrating that prediction targets are geometrically dependent rather than algebraically interchangeable.

flow matchingclean-data predictionlatent spacediffusion modelsvelocity predictionlatent diffusion TransformerFLUX.2 VAEDiTclassifier-free guidanceFID-50KGaussian analysisisotropic target-covariance floor

Abstract

Flow matching with clean-data prediction has shown that regressing the clean point can exploit low-dimensional structure more effectively than predicting an ambient noised quantity. We ask whether this principle remains useful after images are mapped into a learned latent space, where compression has already removed much of the raw pixel variability. We introduce JLT, a 130M latent diffusion Transformer over frozen FLUX.2 VAE codes, and compare clean-latent prediction with a matched velocity-prediction DiT under the same representation, backbone, and training settings. Although the three variables x, epsilon, and v are linearly convertible for a fixed corruption time, a local Gaussian analysis shows that velocity regression inherits an isotropic target-covariance floor and amplifies low-variance latent directions, while clean prediction damps them. On ImageNet 256 x 256, JLT-B/1 obtains FID-50K 2.50 with classifier-free guidance, with a large matched-target gap over velocity prediction. These results suggest that prediction targets in latent diffusion are representation-dependent geometric choices, rather than interchangeable algebraic parameterizations.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
JLT: Clean-Latent Prediction in Latent Diffusion Transformers | TensorX