TensorX
返回文献探索

Paper · arXiv 2406.03184

Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion

Hao Wen, Zehuan Huang, Yaohui Wang, Xinyuan Chen, Yu Qiao, Lu Sheng

20 upvotesJune 5, 2024arXiv 预印本
AI 摘要

A unified framework, Ouroboros3D, combines diffusion-based multi-view image generation and 3D reconstruction through a recursive diffusion process, improving geometric consistency and reducing data bias.

diffusion-based multi-view image generation3D reconstructionrecursive diffusion processself-conditioning mechanismmulti-view denoising3D-aware mapsgeometric consistency

Abstract

Existing single image-to-3D creation methods typically involve a two-stage process, first generating multi-view images, and then using these images for 3D reconstruction. However, training these two stages separately leads to significant data bias in the inference phase, thus affecting the quality of reconstructed results. We introduce a unified 3D generation framework, named Ouroboros3D, which integrates diffusion-based multi-view image generation and 3D reconstruction into a recursive diffusion process. In our framework, these two modules are jointly trained through a self-conditioning mechanism, allowing them to adapt to each other's characteristics for robust inference. During the multi-view denoising process, the multi-view diffusion model uses the 3D-aware maps rendered by the reconstruction module at the previous timestep as additional conditions. The recursive diffusion framework with 3D-aware feedback unites the entire process and improves geometric consistency.Experiments show that our framework outperforms separation of these two stages and existing methods that combine them at the inference phase. Project page: https://costwen.github.io/Ouroboros3D/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号