TensorX
返回文献探索

Paper · arXiv 2305.06351

Reconstructing Animatable Categories from Videos

Gengshan Yang, Chaoyang Wang, N Dinesh Reddy, Deva Ramanan

1 upvotesMay 10, 2023arXiv 预印本
AI 摘要

RAC learns animatable 3D models from monocular videos by optimizing instance-specific skeletons, regularizing latent spaces, and disentangling objects from backgrounds, enabling category-wide models.

differentiable renderingmonocular videoscategory 3D modelsvariation disentanglingmotion disentanglingskeleton optimizationlatent space regularization3D background models

Abstract

Building animatable 3D models is challenging due to the need for 3D scans, laborious registration, and manual rigging, which are difficult to scale to arbitrary categories. Recently, differentiable rendering provides a pathway to obtain high-quality 3D models from monocular videos, but these are limited to rigid categories or single instances. We present RAC that builds category 3D models from monocular videos while disentangling variations over instances and motion over time. Three key ideas are introduced to solve this problem: (1) specializing a skeleton to instances via optimization, (2) a method for latent space regularization that encourages shared structure across a category while maintaining instance details, and (3) using 3D background models to disentangle objects from the background. We show that 3D models of humans, cats, and dogs can be learned from 50-100 internet videos.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Reconstructing Animatable Categories from Videos | TensorX