Contrastive Feature Masking Open-Vocabulary Vision Transformer
Dahun Kim, Anelia Angelova, Weicheng Kuo
CFM-ViT enhances open-vocabulary object detection and image-text retrieval through combined masked and contrastive learning objectives, positional embedding dropout, and region-level representation learning.