TensorX
返回文献探索

Paper · arXiv 2605.16147

Registers Matter for Pixel-Space Diffusion Transformers

Nikita Starodubcev, Ilia Sudakov, Ilya Drobyshevskiy, Artem Babenko, Dmitry Baranchuk

26 upvotesJuly 6, 2026arXiv 预印本
AI 摘要

Register tokens improve diffusion transformers by cleaning high-noise feature maps, and a proposed register guidance technique enhances visual coherence.

Vision Transformersregister tokensDiffusion TransformersDiTspixel-spacelatent-spacefeature mapshigh noise levelsRegister Guidance

Abstract

Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move toward pixel-space training, they become closer in form to ViTs, raising the question of whether register tokens are also useful for Diffusion Transformers (DiTs). In this work, we show that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers. Interestingly, registers are more effective in pixel-space DiTs than in latent-space DiTs. By analyzing intermediate representations, we find that register tokens produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation. We further observe that recent pixel-space DiT architectures implicitly incorporate register-like mechanisms, which may partially account for their strong empirical performance. Motivated by these observations, we propose Register Guidance, a technique that amplifies the contribution of register tokens responsible for improving visual structure and coherence.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Registers Matter for Pixel-Space Diffusion Transformers | TensorX