TensorX
返回文献探索

Paper · arXiv 2401.14398

pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Ege Ozguroglu, Ruoshi Liu, Dídac Surís, Dian Chen, Achal Dave, Pavel Tokmakov, Carl Vondrick

10 upvotesJanuary 25, 2024arXiv 预印本
AI 摘要

Pix2gestalt uses conditional diffusion models, trained on synthetic occluded-object datasets, to achieve zero-shot amodal segmentation better than supervised approaches and enhances object recognition and 3D reconstruction in occlusions.

pix2gestaltzero-shot amodal segmentationdiffusion modelsconditional diffusion modelocclusionssynthetic datasetobject recognition3D reconstruction

Abstract

We introduce pix2gestalt, a framework for zero-shot amodal segmentation, which learns to estimate the shape and appearance of whole objects that are only partially visible behind occlusions. By capitalizing on large-scale diffusion models and transferring their representations to this task, we learn a conditional diffusion model for reconstructing whole objects in challenging zero-shot cases, including examples that break natural and physical priors, such as art. As training data, we use a synthetically curated dataset containing occluded objects paired with their whole counterparts. Experiments show that our approach outperforms supervised baselines on established benchmarks. Our model can furthermore be used to significantly improve the performance of existing object recognition and 3D reconstruction methods in the presence of occlusions.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号