TensorX
返回文献探索

Paper · arXiv 2409.06210

INTRA: Interaction Relationship-aware Weakly Supervised Affordance Grounding

Ji Ha Jang, Hoigi Seo, Se Young Chun

26 upvotesSeptember 10, 2024arXiv 预印本
AI 摘要

INTRA addresses challenges in weakly supervised affordance grounding by using contrastive learning with exocentric images and vision-language embeddings, enhancing flexibility and robustness.

weakly supervised affordance groundingexocentric imagesegocentric image datasetrepresentation learningcontrastive learningvision-language model embeddingstext-conditioned affordance map generationtext synonym augmentationAGD20KIIT-AFFCADUMDdomain scalabilitysynthesized imagesillustrations

Abstract

Affordance denotes the potential interactions inherent in objects. The perception of affordance can enable intelligent agents to navigate and interact with new environments efficiently. Weakly supervised affordance grounding teaches agents the concept of affordance without costly pixel-level annotations, but with exocentric images. Although recent advances in weakly supervised affordance grounding yielded promising results, there remain challenges including the requirement for paired exocentric and egocentric image dataset, and the complexity in grounding diverse affordances for a single object. To address them, we propose INTeraction Relationship-aware weakly supervised Affordance grounding (INTRA). Unlike prior arts, INTRA recasts this problem as representation learning to identify unique features of interactions through contrastive learning with exocentric images only, eliminating the need for paired datasets. Moreover, we leverage vision-language model embeddings for performing affordance grounding flexibly with any text, designing text-conditioned affordance map generation to reflect interaction relationship for contrastive learning and enhancing robustness with our text synonym augmentation. Our method outperformed prior arts on diverse datasets such as AGD20K, IIT-AFF, CAD and UMD. Additionally, experimental results demonstrate that our method has remarkable domain scalability for synthesized images / illustrations and is capable of performing affordance grounding for novel interactions and objects.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号