TensorX
返回文献探索

Paper · arXiv 2402.16936

Disentangled 3D Scene Generation with Layout Learning

Dave Epstein, Ben Poole, Ben Mildenhall, Alexei A. Efros, Aleksander Holynski

11 upvotesFebruary 26, 2024arXiv 预印本
AI 摘要

Unsupervised disentanglement of 3D scenes into component objects using pretrained text-to-image models and joint optimization of NeRFs and layouts.

NeRFsin-distributiontext-to-3Dscene decompositionobject discovery

Abstract

We introduce a method to generate 3D scenes that are disentangled into their component objects. This disentanglement is unsupervised, relying only on the knowledge of a large pretrained text-to-image model. Our key insight is that objects can be discovered by finding parts of a 3D scene that, when rearranged spatially, still produce valid configurations of the same scene. Concretely, our method jointly optimizes multiple NeRFs from scratch - each representing its own object - along with a set of layouts that composite these objects into scenes. We then encourage these composited scenes to be in-distribution according to the image generator. We show that despite its simplicity, our approach successfully generates 3D scenes decomposed into individual objects, enabling new capabilities in text-to-3D content creation. For results and an interactive demo, see our project page at https://dave.ml/layoutlearning/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Disentangled 3D Scene Generation with Layout Learning | TensorX