TensorX
返回文献探索

Paper · arXiv 2404.07199

RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion

Jaidev Shriram, Alex Trevithick, Lingjie Liu, Ravi Ramamoorthi

27 upvotesApril 10, 2024arXiv 预印本
AI 摘要

RealmDreamer generates 3D scenes from text descriptions using 3D Gaussian Splatting, image-conditional diffusion models, and depth diffusion models, achieving high-quality and diverse scene synthesis without requiring video or multi-view data.

3D Gaussian Splattingtext-to-image generatorsimage-conditional diffusion modelsdepth diffusion models3D inpainting tasksharpened samples

Abstract

We introduce RealmDreamer, a technique for generation of general forward-facing 3D scenes from text descriptions. Our technique optimizes a 3D Gaussian Splatting representation to match complex text prompts. We initialize these splats by utilizing the state-of-the-art text-to-image generators, lifting their samples into 3D, and computing the occlusion volume. We then optimize this representation across multiple views as a 3D inpainting task with image-conditional diffusion models. To learn correct geometric structure, we incorporate a depth diffusion model by conditioning on the samples from the inpainting model, giving rich geometric structure. Finally, we finetune the model using sharpened samples from image generators. Notably, our technique does not require video or multi-view data and can synthesize a variety of high-quality 3D scenes in different styles, consisting of multiple objects. Its generality additionally allows 3D synthesis from a single image.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion | TensorX