TensorX
返回文献探索

Paper · arXiv 2402.07207

GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

Xiaoyu Zhou, Xingjian Ran, Yajiao Xiong, Jinlin He, Zhiwei Lin, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang

10 upvotesFebruary 11, 2024arXiv 预印本
AI 摘要

GALA3D is a generative framework for 3D content that uses large language models and conditioned diffusion for layout-guided, compositional scene generation, ensuring high fidelity and realistic object interactions.

large language modelsLLMslayout-guided3D Gaussianobject-scene compositional optimizationconditioned diffusion.scene-level 3D content generationcontrollable editinghigh fidelityobject-level entities

Abstract

We present GALA3D, generative 3D GAussians with LAyout-guided control, for effective compositional text-to-3D generation. We first utilize large language models (LLMs) to generate the initial layout and introduce a layout-guided 3D Gaussian representation for 3D content generation with adaptive geometric constraints. We then propose an object-scene compositional optimization mechanism with conditioned diffusion to collaboratively generate realistic 3D scenes with consistent geometry, texture, scale, and accurate interactions among multiple objects while simultaneously adjusting the coarse layout priors extracted from the LLMs to align with the generated scene. Experiments show that GALA3D is a user-friendly, end-to-end framework for state-of-the-art scene-level 3D content generation and controllable editing while ensuring the high fidelity of object-level entities within the scene. Source codes and models will be available at https://gala3d.github.io/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting | TensorX