TensorX
返回文献探索

Paper · arXiv 2404.16030

MoDE: CLIP Data Experts via Clustering

Jiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li, Luke Zettlemoyer, Shih-Fu Chang, Wen-Tau Yih, Hu Xu

13 upvotesApril 24, 2024arXiv 预印本
AI 摘要

Mixture of Data Experts (MoDE) enhances CLIP's performance by clustering data into experts, reducing noise and improving zero-shot classification with lower training costs.

contrastive language-image pretrainingCLIPsupervisiondata expertsclusteringfalse negative noisesensembletask metadatacluster conditionsontologyfine-grained cluster centerscoarse-grained levelzero-shot image classificationtraining costasynchronous training

Abstract

The success of contrastive language-image pretraining (CLIP) relies on the supervision from the pairing between images and captions, which tends to be noisy in web-crawled data. We present Mixture of Data Experts (MoDE) and learn a system of CLIP data experts via clustering. Each data expert is trained on one data cluster, being less sensitive to false negative noises in other clusters. At inference time, we ensemble their outputs by applying weights determined through the correlation between task metadata and cluster conditions. To estimate the correlation precisely, the samples in one cluster should be semantically similar, but the number of data experts should still be reasonable for training and inference. As such, we consider the ontology in human language and propose to use fine-grained cluster centers to represent each data expert at a coarse-grained level. Experimental studies show that four CLIP data experts on ViT-B/16 outperform the ViT-L/14 by OpenAI CLIP and OpenCLIP on zero-shot image classification but with less (<35\%) training cost. Meanwhile, MoDE can train all data expert asynchronously and can flexibly include new data experts. The code is available at https://github.com/facebookresearch/MetaCLIP/tree/main/mode.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MoDE: CLIP Data Experts via Clustering | TensorX