TensorX
返回文献探索

Paper · arXiv 2305.18766

HiFA: High-fidelity Text-to-3D with Advanced Diffusion Guidance

Joseph Zhu, Peiye Zhuang

7 upvotesMay 30, 2023arXiv 预印本
AI 摘要

Reformulating optimization loss with diffusion prior and introducing depth supervision improve text-to-3D synthesis quality and multi-view consistency.

diffusion modelsNeural Radiance Fields (NeRFs)diffusion priordepth supervisiondensity field

Abstract

Automatic text-to-3D synthesis has achieved remarkable advancements through the optimization of 3D models. Existing methods commonly rely on pre-trained text-to-image generative models, such as diffusion models, providing scores for 2D renderings of Neural Radiance Fields (NeRFs) and being utilized for optimizing NeRFs. However, these methods often encounter artifacts and inconsistencies across multiple views due to their limited understanding of 3D geometry. To address these limitations, we propose a reformulation of the optimization loss using the diffusion prior. Furthermore, we introduce a novel training approach that unlocks the potential of the diffusion prior. To improve 3D geometry representation, we apply auxiliary depth supervision for NeRF-rendered images and regularize the density field of NeRFs. Extensive experiments demonstrate the superiority of our method over prior works, resulting in advanced photo-realism and improved multi-view consistency.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
HiFA: High-fidelity Text-to-3D with Advanced Diffusion Guidance | TensorX