TensorX
返回文献探索

Paper · arXiv 2310.13355

SILC: Improving Vision Language Pretraining with Self-Distillation

Muhammad Ferjad Naeem, Yongqin Xian, Xiaohua Zhai, Lukas Hoyer, Luc Van Gool, Federico Tombari

9 upvotesOctober 20, 2023arXiv 预印本
AI 摘要

Adding local-to-global correspondence learning via self-distillation to contrastive pre-training improves model performance across various computer vision tasks, particularly segmentation, and sets new benchmarks in zero-shot classification, few-shot classification, and retrieval.

contrastive objectivelocal-to-global correspondenceself-distillationexponential moving average (EMA)model performancecomputer vision tasksclassificationretrievalsegmentationzero-shot classificationfew-shot classificationzero-shot segmentationopen vocabulary segmentation

Abstract

Image-Text pretraining on web-scale image caption dataset has become the default recipe for open vocabulary classification and retrieval models thanks to the success of CLIP and its variants. Several works have also used CLIP features for dense prediction tasks and have shown the emergence of open-set abilities. However, the contrastive objective only focuses on image-text alignment and does not incentivise image feature learning for dense prediction tasks. In this work, we propose the simple addition of local-to-global correspondence learning by self-distillation as an additional objective for contrastive pre-training to propose SILC. We show that distilling local image features from an exponential moving average (EMA) teacher model significantly improves model performance on several computer vision tasks including classification, retrieval, and especially segmentation. We further show that SILC scales better with the same training duration compared to the baselines. Our model SILC sets a new state of the art for zero-shot classification, few shot classification, image and text retrieval, zero-shot segmentation, and open vocabulary segmentation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SILC: Improving Vision Language Pretraining with Self-Distillation | TensorX