TensorX
返回文献探索

Paper · arXiv 2312.04875

MVDD: Multi-View Depth Diffusion Models

Zhen Wang, Qiangeng Xu, Feitong Tan, Menglei Chai, Shichen Liu, Rohit Pandey, Sean Fanello, Achuta Kadambi, Yinda Zhang

10 upvotesDecember 8, 2023arXiv 预印本
AI 摘要

MVDD is a denoising diffusion model using multi-view depth representations to generate high-quality 3D shape point clouds and meshes, achieving state-of-the-art results in 3D shape generation and depth completion.

denoising diffusion modelsmulti-view depthepipolar line segment attentiondepth fusion modulesurface reconstructiondepth completionGAN inversion3D shape generation3D prior

Abstract

Denoising diffusion models have demonstrated outstanding results in 2D image generation, yet it remains a challenge to replicate its success in 3D shape generation. In this paper, we propose leveraging multi-view depth, which represents complex 3D shapes in a 2D data format that is easy to denoise. We pair this representation with a diffusion model, MVDD, that is capable of generating high-quality dense point clouds with 20K+ points with fine-grained details. To enforce 3D consistency in multi-view depth, we introduce an epipolar line segment attention that conditions the denoising step for a view on its neighboring views. Additionally, a depth fusion module is incorporated into diffusion steps to further ensure the alignment of depth maps. When augmented with surface reconstruction, MVDD can also produce high-quality 3D meshes. Furthermore, MVDD stands out in other tasks such as depth completion, and can serve as a 3D prior, significantly boosting many downstream tasks, such as GAN inversion. State-of-the-art results from extensive experiments demonstrate MVDD's excellent ability in 3D shape generation, depth completion, and its potential as a 3D prior for downstream tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MVDD: Multi-View Depth Diffusion Models | TensorX