TensorX
返回文献探索

Paper · arXiv 2306.04619

ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image Collections

Chun-Han Yao, Amit Raj, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang, Varun Jampani

4 upvotesJune 7, 2023arXiv 预印本
AI 摘要

ARTIC3D reconstructs 3D articulated shapes from complex monocular images using a skeleton-based model, 2D diffusion priors, and diffusion-guided optimization, achieving higher fidelity and more realistic animations.

skeleton-based surface representation2D diffusion priorsStable Diffusiondiffusion-guided 3D optimizationimage-level gradientsdiffusion modelsrigid part transformations

Abstract

Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc. We propose ARTIC3D, a self-supervised framework to reconstruct per-instance 3D shapes from a sparse image collection in-the-wild. Specifically, ARTIC3D is built upon a skeleton-based surface representation and is further guided by 2D diffusion priors from Stable Diffusion. First, we enhance the input images with occlusions/truncation via 2D diffusion to obtain cleaner mask estimates and semantic features. Second, we perform diffusion-guided 3D optimization to estimate shape and texture that are of high-fidelity and faithful to input images. We also propose a novel technique to calculate more stable image-level gradients via diffusion models compared to existing alternatives. Finally, we produce realistic animations by fine-tuning the rendered shape and texture under rigid part transformations. Extensive evaluations on multiple existing datasets as well as newly introduced noisy web image collections with occlusions and truncation demonstrate that ARTIC3D outputs are more robust to noisy images, higher quality in terms of shape and texture details, and more realistic when animated. Project page: https://chhankyao.github.io/artic3d/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号