TensorX
返回文献探索

Paper · arXiv 2407.13764

Shape of Motion: 4D Reconstruction from a Single Video

Qianqian Wang, Vickie Ye, Hang Gao, Jake Austin, Zhengqi Li, Angjoo Kanazawa

20 upvotesJuly 18, 2024arXiv 预印本
AI 摘要

A monocular dynamic reconstruction method using SE3 motion bases and data-driven priors achieves state-of-the-art performance in 3D/2D motion estimation and novel view synthesis.

monocular dynamic reconstructionSE3 motion bases3D motionlow-dimensional structuredata-driven priorsmonocular depth mapslong-range 2D tracksmotion estimationnovel view synthesis

Abstract

Monocular dynamic reconstruction is a challenging and long-standing vision problem due to the highly ill-posed nature of the task. Existing approaches are limited in that they either depend on templates, are effective only in quasi-static scenes, or fail to model 3D motion explicitly. In this work, we introduce a method capable of reconstructing generic dynamic scenes, featuring explicit, full-sequence-long 3D motion, from casually captured monocular videos. We tackle the under-constrained nature of the problem with two key insights: First, we exploit the low-dimensional structure of 3D motion by representing scene motion with a compact set of SE3 motion bases. Each point's motion is expressed as a linear combination of these bases, facilitating soft decomposition of the scene into multiple rigidly-moving groups. Second, we utilize a comprehensive set of data-driven priors, including monocular depth maps and long-range 2D tracks, and devise a method to effectively consolidate these noisy supervisory signals, resulting in a globally consistent representation of the dynamic scene. Experiments show that our method achieves state-of-the-art performance for both long-range 3D/2D motion estimation and novel view synthesis on dynamic scenes. Project Page: https://shape-of-motion.github.io/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Shape of Motion: 4D Reconstruction from a Single Video | TensorX