TensorX
返回文献探索

Paper · arXiv 2311.06214

Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model

Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, Sai Bi

33 upvotesNovember 10, 2023arXiv 预印本
AI 摘要

Instant3D rapidly generates high-quality 3D assets from text using a two-stage approach involving a fine-tuned 2D diffusion model and a transformer-based reconstructor.

text-to-3Ddiffusion modelsscore distillation-based optimizationJanus problemsfeed-forward methodssparse set of views2D text-to-image diffusion modelNeRFtransformer-based sparse-view reconstructor

Abstract

Text-to-3D with diffusion models have achieved remarkable progress in recent years. However, existing methods either rely on score distillation-based optimization which suffer from slow inference, low diversity and Janus problems, or are feed-forward methods that generate low quality results due to the scarcity of 3D training data. In this paper, we propose Instant3D, a novel method that generates high-quality and diverse 3D assets from text prompts in a feed-forward manner. We adopt a two-stage paradigm, which first generates a sparse set of four structured and consistent views from text in one shot with a fine-tuned 2D text-to-image diffusion model, and then directly regresses the NeRF from the generated images with a novel transformer-based sparse-view reconstructor. Through extensive experiments, we demonstrate that our method can generate high-quality, diverse and Janus-free 3D assets within 20 seconds, which is two order of magnitude faster than previous optimization-based methods that can take 1 to 10 hours. Our project webpage: https://jiahao.ai/instant3d/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model | TensorX