TensorX
返回文献探索

Paper · arXiv 2507.13347

π^3: Scalable Permutation-Equivariant Visual Geometry Learning

Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chunhua Shen, Tong He

67 upvotesJuly 17, 2025arXiv 预印本
AI 摘要

A permutation-equivariant neural network, $\pi^3$, reconstructs visual geometry without a fixed reference view, achieving state-of-the-art performance in camera pose estimation, depth estimation, and point map reconstruction.

feed-forward neural networkpermutation-equivariant architectureaffine-invariantscale-invariantcamera pose estimationmonocular depth estimationvideo depth estimationdense point map reconstruction

Abstract

We introduce pi^3, a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Previous methods often anchor their reconstructions to a designated viewpoint, an inductive bias that can lead to instability and failures if the reference is suboptimal. In contrast, pi^3 employs a fully permutation-equivariant architecture to predict affine-invariant camera poses and scale-invariant local point maps without any reference frames. This design makes our model inherently robust to input ordering and highly scalable. These advantages enable our simple and bias-free approach to achieve state-of-the-art performance on a wide range of tasks, including camera pose estimation, monocular/video depth estimation, and dense point map reconstruction. Code and models are publicly available.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
π^3: Scalable Permutation-Equivariant Visual Geometry Learning | TensorX