TensorX
返回文献探索

Paper · arXiv 2311.04400

LRM: Large Reconstruction Model for Single Image to 3D

Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, Hao Tan

52 upvotesNovember 8, 2023arXiv 预印本
AI 摘要

A Large Reconstruction Model using a transformer-based architecture predicts 3D neural radiance fields from single images using massive multi-view training data.

Large Reconstruction Modeltransformer-based architectureneural radiance fieldShapeNetObjaverseMVImgNetend-to-end training

Abstract

We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. We train our model in an end-to-end manner on massive multi-view data containing around 1 million objects, including both synthetic renderings from Objaverse and real captures from MVImgNet. This combination of a high-capacity model and large-scale training data empowers our model to be highly generalizable and produce high-quality 3D reconstructions from various testing inputs including real-world in-the-wild captures and images from generative models. Video demos and interactable 3D meshes can be found on this website: https://yiconghong.me/LRM/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LRM: Large Reconstruction Model for Single Image to 3D | TensorX