TensorX
返回文献探索

Paper · arXiv 2411.08033

GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation

Yushi Lan, Shangchen Zhou, Zhaoyang Lyu, Fangzhou Hong, Shuai Yang, Bo Dai, Xingang Pan, Chen Change Loy

25 upvotesNovember 12, 2024arXiv 预印本
AI 摘要

A novel 3D generation framework using a VAE and cascaded latent diffusion model in a Point Cloud-structured Latent space achieves high-quality 3D generation with multi-modal inputs and disentangled shape-texture editing.

Variational Autoencoderlatent spaceRGB-Dmulti-viewlatent diffusion modelshape-texture disentanglementPoint Cloud-structured Latent spaceGaussianAnything3D-aware editing

Abstract

While 3D content generation has advanced significantly, existing methods still face challenges with input formats, latent space design, and output representations. This paper introduces a novel 3D generation framework that addresses these challenges, offering scalable, high-quality 3D generation with an interactive Point Cloud-structured Latent space. Our framework employs a Variational Autoencoder (VAE) with multi-view posed RGB-D(epth)-N(ormal) renderings as input, using a unique latent space design that preserves 3D shape information, and incorporates a cascaded latent diffusion model for improved shape-texture disentanglement. The proposed method, GaussianAnything, supports multi-modal conditional 3D generation, allowing for point cloud, caption, and single/multi-view image inputs. Notably, the newly proposed latent space naturally enables geometry-texture disentanglement, thus allowing 3D-aware editing. Experimental results demonstrate the effectiveness of our approach on multiple datasets, outperforming existing methods in both text- and image-conditioned 3D generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号