TensorX
返回文献探索

Paper · arXiv 2401.07727

HexaGen3D: StableDiffusion is just one step away from Fast and Diverse Text-to-3D Generation

Antoine Mercier, Ramin Nakhli, Mahesh Reddy, Rajeev Yasarla, Hong Cai, Fatih Porikli, Guillaume Berger

11 upvotesJanuary 15, 2024arXiv 预印本
AI 摘要

HexaGen3D fine-tunes a 2D diffusion model to generate high-quality 3D assets from textual prompts using 6 orthographic projections and a latent triplane, offering better quality-to-latency trade-offs than existing methods.

diffusion modelstext-to-imageorthographic projectionslatent triplanetextured meshper-sample optimization

Abstract

Despite the latest remarkable advances in generative modeling, efficient generation of high-quality 3D assets from textual prompts remains a difficult task. A key challenge lies in data scarcity: the most extensive 3D datasets encompass merely millions of assets, while their 2D counterparts contain billions of text-image pairs. To address this, we propose a novel approach which harnesses the power of large, pretrained 2D diffusion models. More specifically, our approach, HexaGen3D, fine-tunes a pretrained text-to-image model to jointly predict 6 orthographic projections and the corresponding latent triplane. We then decode these latents to generate a textured mesh. HexaGen3D does not require per-sample optimization, and can infer high-quality and diverse objects from textual prompts in 7 seconds, offering significantly better quality-to-latency trade-offs when comparing to existing approaches. Furthermore, HexaGen3D demonstrates strong generalization to new objects or compositions.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
HexaGen3D: StableDiffusion is just one step away from Fast and Diverse Text-to-3D Generation | TensorX