TensorX
返回文献探索

Paper · arXiv 2311.09217

DMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction Model

Yinghao Xu, Hao Tan, Fujun Luan, Sai Bi, Peng Wang, Jiahao Li, Zifan Shi, Kalyan Sunkavalli, Gordon Wetzstein, Zexiang Xu, Kai Zhang

22 upvotesNovember 15, 2023arXiv 预印本
AI 摘要

DMV3D, a transformer-based 3D generation model using a triplane NeRF representation, achieves state-of-the-art results in single-image and text-to-3D reconstruction tasks within 30 seconds.

DMV3Dtransformer-based3D generationdenoisemulti-view diffusiontriplane NeRFNeRFimage reconstruction lossessingle-image reconstructionprobabilistic modelingtext-to-3D generation

Abstract

We propose DMV3D, a novel 3D generation approach that uses a transformer-based 3D large reconstruction model to denoise multi-view diffusion. Our reconstruction model incorporates a triplane NeRF representation and can denoise noisy multi-view images via NeRF reconstruction and rendering, achieving single-stage 3D generation in sim30s on single A100 GPU. We train DMV3D on large-scale multi-view image datasets of highly diverse objects using only image reconstruction losses, without accessing 3D assets. We demonstrate state-of-the-art results for the single-image reconstruction problem where probabilistic modeling of unseen object parts is required for generating diverse reconstructions with sharp textures. We also show high-quality text-to-3D generation results outperforming previous 3D diffusion models. Our project website is at: https://justimyhxu.github.io/projects/dmv3d/ .

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
DMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction Model | TensorX