TensorX
返回文献探索

Paper · arXiv 2405.16822

Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels

Yikai Wang, Xinzhou Wang, Zilong Chen, Zhengyi Wang, Fuchun Sun, Jun Zhu

13 upvotesMay 27, 2024arXiv 预印本
AI 摘要

Vidu4D, a novel reconstruction model using Dynamic Gaussian Surfels, accurately reconstructs 4D representations from single generated videos, addressing non-rigidity and frame distortion for high-fidelity virtual contents.

reconstruction modelDynamic Gaussian SurfelsGaussian surfelstime-varying warping functionswarped-state geometric regularizationcontinuous warping fieldsrotation parametersscaling parameterstexture flickeringhigh-fidelity text-to-4D generation

Abstract

Video generative models are receiving particular attention given their ability to generate realistic and imaginative frames. Besides, these models are also observed to exhibit strong 3D consistency, significantly enhancing their potential to act as world simulators. In this work, we present Vidu4D, a novel reconstruction model that excels in accurately reconstructing 4D (i.e., sequential 3D) representations from single generated videos, addressing challenges associated with non-rigidity and frame distortion. This capability is pivotal for creating high-fidelity virtual contents that maintain both spatial and temporal coherence. At the core of Vidu4D is our proposed Dynamic Gaussian Surfels (DGS) technique. DGS optimizes time-varying warping functions to transform Gaussian surfels (surface elements) from a static state to a dynamically warped state. This transformation enables a precise depiction of motion and deformation over time. To preserve the structural integrity of surface-aligned Gaussian surfels, we design the warped-state geometric regularization based on continuous warping fields for estimating normals. Additionally, we learn refinements on rotation and scaling parameters of Gaussian surfels, which greatly alleviates texture flickering during the warping process and enhances the capture of fine-grained appearance details. Vidu4D also contains a novel initialization state that provides a proper start for the warping fields in DGS. Equipping Vidu4D with an existing video generative model, the overall framework demonstrates high-fidelity text-to-4D generation in both appearance and geometry.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号