TensorX
返回文献探索

Paper · arXiv 2312.14232

Parrot Captions Teach CLIP to Spot Text

Yiqi Lin, Conghui He, Alex Jinpeng Wang, Bin Wang, Weijia Li, Mike Zheng Shou

11 upvotesDecember 21, 2023arXiv 预印本
AI 摘要

CLIP models exhibit a text spotting bias in vision-language applications due to excessive reliance on visual text within images in datasets like LAION-2B.

CLIPvision-language applicationstext spotting biasLAION-2Bvisual textimage-text similarityvisual-language representation learningCLIP-like modelsimage-text dataset curation

Abstract

Despite CLIP being the foundation model in numerous vision-language applications, the CLIP suffers from a severe text spotting bias. Such bias causes CLIP models to `Parrot' the visual text embedded within images while disregarding the authentic visual semantics. We uncover that in the most popular image-text dataset LAION-2B, the captions also densely parrot (spell) the text embedded in images. Our analysis shows that around 50\% of images are embedded with visual text content, and 90\% of their captions more or less parrot the visual text. Based on such observation, we thoroughly inspect the different release d versions of CLIP models and verify that the visual text is the dominant factor in measuring the LAION-style image-text similarity for these models. To examine whether these parrot captions shape the text spotting bias, we train a series of CLIP models with LAION subsets curated by different parrot-caption-oriented criteria. We show that training with parrot captions easily shapes such bias but harms the expected visual-language representation learning in CLIP models. This suggests that it is urgent to revisit either the design of CLIP-like models or the existing image-text dataset curation pipeline built on CLIP score filtering.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Parrot Captions Teach CLIP to Spot Text | TensorX