TensorX
返回文献探索

Paper · arXiv 2410.02746

Contrastive Localized Language-Image Pre-Training

Hong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang, Marcin Eichner, Keen You, Meng Cao, Bowen Zhang, Yinfei Yang, Zhe Gan

36 upvotesOctober 3, 2024arXiv 预印本
AI 摘要

CLOC enhances CLIP's localization capabilities by introducing region-text contrastive loss and promptable embeddings, improving regional image representation for multimodal large language models.

Contrastive Language-Image Pre-training (CLIP)multimodal large language models (MLLMs)region-text contrastive losspromptable embeddingsregion representationsvisually-enriched captioning frameworkregional embeddingsimage region recognitionretrieval tasks

Abstract

Contrastive Language-Image Pre-training (CLIP) has been a celebrated method for training vision encoders to generate image/text representations facilitating various applications. Recently, CLIP has been widely adopted as the vision backbone of multimodal large language models (MLLMs) to connect image inputs for language interactions. The success of CLIP as a vision-language foundation model relies on aligning web-crawled noisy text annotations at image levels. Nevertheless, such criteria may become insufficient for downstream tasks in need of fine-grained vision representations, especially when region-level understanding is demanding for MLLMs. In this paper, we improve the localization capability of CLIP with several advances. We propose a pre-training method called Contrastive Localized Language-Image Pre-training (CLOC) by complementing CLIP with region-text contrastive loss and modules. We formulate a new concept, promptable embeddings, of which the encoder produces image embeddings easy to transform into region representations given spatial hints. To support large-scale pre-training, we design a visually-enriched and spatially-localized captioning framework to effectively generate region-text pseudo-labels at scale. By scaling up to billions of annotated images, CLOC enables high-quality regional embeddings for image region recognition and retrieval tasks, and can be a drop-in replacement of CLIP to enhance MLLMs, especially on referring and grounding tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Contrastive Localized Language-Image Pre-Training | TensorX