TensorX
返回文献探索

Paper · arXiv 2407.07860

Controlling Space and Time with Diffusion Models

Daniel Watson, Saurabh Saxena, Lala Li, Andrea Tagliasacchi, David J. Fleet

16 upvotesJuly 10, 2024arXiv 预印本
AI 摘要

Cascaded diffusion model for 4D novel view synthesis with joint 3D, 4D, and video training, enabling metric scale camera control and state-of-the-art fidelity with temporal dynamics handling.

cascaded diffusion model4D novel view synthesis4D training datajoint training3D trainingvideo trainingcamera posestimestampsmonocular metric depth estimatorsmetric scale camera controlevaluation metricsstate-of-the-artpose controltemporal dynamicspanorama stitchingpose-conditioned video to video translation

Abstract

We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), conditioned on one or more images of a general scene, and a set of camera poses and timestamps. To overcome challenges due to limited availability of 4D training data, we advocate joint training on 3D (with camera pose), 4D (pose+time) and video (time but no pose) data and propose a new architecture that enables the same. We further advocate the calibration of SfM posed data using monocular metric depth estimators for metric scale camera control. For model evaluation, we introduce new metrics to enrich and overcome shortcomings of current evaluation schemes, demonstrating state-of-the-art results in both fidelity and pose control compared to existing diffusion models for 3D NVS, while at the same time adding the ability to handle temporal dynamics. 4DiM is also used for improved panorama stitching, pose-conditioned video to video translation, and several other tasks. For an overview see https://4d-diffusion.github.io

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号