TensorX
返回文献探索

Paper · arXiv 2305.03689

COLA: How to adapt vision-language models to Compose Objects Localized with Attributes?

Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan A. Plummer, Ranjay Krishna, Kate Saenko

5 upvotesMay 5, 2023arXiv 预印本
AI 摘要

A novel text-to-image retrieval benchmark (Cola) is used to assess and enhance compositional reasoning in vision-language models through fine-tuning strategies.

vision-language modelscompositional reasoningtext-to-image retrievalColamulti-modal transformerpretrainingfine-tuning strategiesCLIPFLAVAmulti-modal adapter

Abstract

Compositional reasoning is a hallmark of human visual intelligence; yet despite the size of large vision-language models, they struggle to represent simple compositions by combining objects with their attributes. To measure this lack of compositional capability, we design Cola, a text-to-image retrieval benchmark to Compose Objects Localized with Attributes. Using Cola as a testbed, we explore modeling designs to adapt pre-trained vision-language models to reason compositionally about multiple attributes attached to multiple objects. We explore 6 finetuning strategies on 2 seminal vision-language models, using 3 finetuning datasets and 2 test benchmarks (Cola and CREPE). Surprisingly, our optimal finetuning strategy improves a 151M parameter CLIP, which disjointly encodes image and language during pretraining, to perform as well as a 241M parameter FLAVA, which uses a multi-modal transformer encoder during pretraining to attend over both vision and language modalities. This optimal finetuning strategy is a lightweight multi-modal adapter that jointly attends over both image and language features generated by the pretrained model. We show this works better than common strategies such as prompt/fine-tuning, or tuning a comparable number of unimodal layers.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
COLA: How to adapt vision-language models to Compose Objects Localized with Attributes? | TensorX