TensorX
返回文献探索

Paper · arXiv 2412.01506

Structured 3D Latents for Scalable and Versatile 3D Generation

Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, Jiaolong Yang

92 upvotesDecember 2, 2024arXiv 预印本
AI 摘要

A 3D generation method using a unified SLAT representation and rectified flow transformers achieves high-quality results across different formats and conditions.

Structured LATent (SLAT)Radiance Fields3D Gaussiansmeshessparsely-populated 3D griddense multiview visual featuresrectified flow transformers

Abstract

We introduce a novel 3D generation method for versatile and high-quality 3D asset creation. The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D Gaussians, and meshes. This is achieved by integrating a sparsely-populated 3D grid with dense multiview visual features extracted from a powerful vision foundation model, comprehensively capturing both structural (geometry) and textural (appearance) information while maintaining flexibility during decoding. We employ rectified flow transformers tailored for SLAT as our 3D generation models and train models with up to 2 billion parameters on a large 3D asset dataset of 500K diverse objects. Our model generates high-quality results with text or image conditions, significantly surpassing existing methods, including recent ones at similar scales. We showcase flexible output format selection and local 3D editing capabilities which were not offered by previous models. Code, model, and data will be released.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Structured 3D Latents for Scalable and Versatile 3D Generation | TensorX