TensorX
返回文献探索

Paper · arXiv 2607.02515

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

Haofei Xu, Rundi Wu, Philipp Henzler, Nikolai Kalischek, Michael Oechsle, Fabian Manhardt, Marc Pollefeys, Andreas Geiger, Federico Tombari, Michael Niemeyer

22 upvotesJuly 2, 2026arXiv 预印本
AI 摘要

A minimalist pixel-space diffusion transformer using plain ViT architecture directly processes 3D point map patches conditioned on image tokens from DINOv3, outperforming complex latent-based models while maintaining simplicity and robustness in ambiguous regions.

Diffusion TransformerViT3D point map patchesDINOv3pixel-spacelatent diffusion modelshybrid architecturesloss functionsgeometric structuretransparent objects

Abstract

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulations are unnecessary. We introduce a minimalist pixel-space Diffusion Transformer, built on a plain ViT, that operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3. Unlike existing latent diffusion approaches, we train our diffusion backbone entirely from scratch, eliminating the need for point map tokenizers. Despite its simplicity, our approach surpasses complex latent-based diffusion models while remaining significantly simpler than hybrid alternatives. Notably, it produces sharper geometric structure and is more robust in highly ambiguous regions, such as transparent objects.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation | TensorX