TensorX
返回文献探索

Paper · arXiv 2408.16061

3D Reconstruction with Spatial Memory

Hengyi Wang, Lourdes Agapito

15 upvotesAugust 28, 2024arXiv 预印本
AI 摘要

Spann3R is a transformer-based method for dense 3D reconstruction from image collections, using an external spatial memory to predict global pointmaps and demonstrating real-time performance and generalization.

transformer-based architecturepointmapsDUSt3R paradigmglobal coordinate systemspatial memorypre-trained weights

Abstract

We present Spann3R, a novel approach for dense 3D reconstruction from ordered or unordered image collections. Built on the DUSt3R paradigm, Spann3R uses a transformer-based architecture to directly regress pointmaps from images without any prior knowledge of the scene or camera parameters. Unlike DUSt3R, which predicts per image-pair pointmaps each expressed in its local coordinate frame, Spann3R can predict per-image pointmaps expressed in a global coordinate system, thus eliminating the need for optimization-based global alignment. The key idea of Spann3R is to manage an external spatial memory that learns to keep track of all previous relevant 3D information. Spann3R then queries this spatial memory to predict the 3D structure of the next frame in a global coordinate system. Taking advantage of DUSt3R's pre-trained weights, and further fine-tuning on a subset of datasets, Spann3R shows competitive performance and generalization ability on various unseen datasets and can process ordered image collections in real time. Project page: https://hengyiwang.github.io/projects/spanner

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
3D Reconstruction with Spatial Memory | TensorX