TensorX
返回文献探索

Paper · arXiv 2507.12462

SpatialTrackerV2: 3D Point Tracking Made Easy

Yuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev, Yuri Makarov, Bingyi Kang, Xing Zhu, Hujun Bao, Yujun Shen, Xiaowei Zhou

19 upvotesJuly 16, 2025arXiv 预印本
AI 摘要

SpatialTrackerV2 is a feed-forward 3D point tracking method for monocular videos that integrates scene geometry, camera ego-motion, and object motion into a unified, differentiable architecture, achieving high accuracy and speed.

feed-forward3D point trackingmonocular videosscene geometrycamera ego-motionpixel-wise object motionfully differentiableend-to-end architecturesynthetic sequencesposed RGB-D videosunlabeled in-the-wild footagedynamic 3D reconstruction

Abstract

We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50times faster.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号