TensorX
返回文献探索

Paper · arXiv 2308.03793

ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain Adaptation

Hu. Xuefeng, Zhang. Ke, Xia. Lu, Chen. Albert, Luo. Jiajia, Sun. Yuyin, Wang. Ken, Qiao. Nan, Zeng. Xiao, Sun. Min, Kuo. Cheng-Hao, Nevatia. Ram

11 upvotesAugust 4, 2023arXiv 预印本
AI 摘要

ReCLIP, a source-free domain adaptation method for vision-language models, reduces the average error rate of CLIP by learning a projection space and deploying cross-modality self-training with pseudo labels.

vision-language modelsCLIPzero-shot classificationdomain gapscross-modality misalignmentsource-free domain adaptationprojection spacepseudo labelscross-modality self-trainingvisual encoderstext encoders

Abstract

Large-scale Pre-Training Vision-Language Model such as CLIP has demonstrated outstanding performance in zero-shot classification, e.g. achieving 76.3% top-1 accuracy on ImageNet without seeing any example, which leads to potential benefits to many tasks that have no labeled data. However, while applying CLIP to a downstream target domain, the presence of visual and text domain gaps and cross-modality misalignment can greatly impact the model performance. To address such challenges, we propose ReCLIP, the first source-free domain adaptation method for vision-language models, which does not require any source data or target labeled data. ReCLIP first learns a projection space to mitigate the misaligned visual-text embeddings and learns pseudo labels, and then deploys cross-modality self-training with the pseudo labels, to update visual and text encoders, refine labels and reduce domain gaps and misalignments iteratively. With extensive experiments, we demonstrate ReCLIP reduces the average error rate of CLIP from 30.17% to 25.06% on 22 image classification benchmarks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain Adaptation | TensorX