TensorX
返回文献探索

Paper · arXiv 2306.03881

Emergent Correspondence from Image Diffusion

Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, Bharath Hariharan

6 upvotesJune 6, 2023arXiv 预印本
AI 摘要

Image diffusion models extract implicit features (DIFT) that outperform supervised and weakly-supervised methods in establishing semantic, geometric, and temporal correspondences between images.

diffusion modelsDIffusion FeaTures (DIFT)semantic correspondencegeometric correspondencetemporal correspondenceStable DiffusionDINOOpenCLIPSPair-71k benchmark

Abstract

Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We propose a simple strategy to extract this implicit knowledge out of diffusion networks as image features, namely DIffusion FeaTures (DIFT), and use them to establish correspondences between real images. Without any additional fine-tuning or supervision on the task-specific data or annotations, DIFT is able to outperform both weakly-supervised methods and competitive off-the-shelf features in identifying semantic, geometric, and temporal correspondences. Particularly for semantic correspondence, DIFT from Stable Diffusion is able to outperform DINO and OpenCLIP by 19 and 14 accuracy points respectively on the challenging SPair-71k benchmark. It even outperforms the state-of-the-art supervised methods on 9 out of 18 categories while remaining on par for the overall performance. Project page: https://diffusionfeatures.github.io

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Emergent Correspondence from Image Diffusion | TensorX