TensorX
返回文献探索

Paper · arXiv 2405.20204

Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Andreas Koukounas, Georgios Mastrapas, Michael Günther, Bo Wang, Scott Martens, Isabelle Mohr, Saba Sturua, Mohammad Kalim Akram, Joan Fontanals Martínez, Saahil Ognawala, Susana Guzman, Maximilian Werk, Nan Wang, Han Xiao

37 upvotesMay 30, 2024arXiv 预印本
AI 摘要

A novel multi-task contrastive training method improves CLIP model performance on both text-image and text-text retrieval tasks.

Contrastive Language-Image PretrainingCLIPembedding spacemultimodal information retrievaltext-only tasksmulti-task contrastive trainingjina-clip-v1

Abstract

Contrastive Language-Image Pretraining (CLIP) is widely used to train models to align images and texts in a common embedding space by mapping them to fixed-sized vectors. These models are key to multimodal information retrieval and related tasks. However, CLIP models generally underperform in text-only tasks compared to specialized text models. This creates inefficiencies for information retrieval systems that keep separate embeddings and models for text-only and multimodal tasks. We propose a novel, multi-task contrastive training method to address this issue, which we use to train the jina-clip-v1 model to achieve the state-of-the-art performance on both text-image and text-text retrieval tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号