TensorX
返回文献探索

Paper · arXiv 2305.07011

Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers

Dahun Kim, Anelia Angelova, Weicheng Kuo

6 upvotesMay 11, 2023arXiv 预印本
AI 摘要

Region-aware Open-vocabulary Vision Transformers improve open-vocabulary object detection and image-text retrieval by using focal loss and novel object proposals in contrastive learning.

Region-aware Open-vocabulary Vision TransformersRO-ViTcontrastive image-text pretrainingpositional embeddingsfocal lossopen-vocabulary object detectionLVISCOCOimage-text retrievalzero-shot transferAP_rimage-level representationnovel object proposals

Abstract

We present Region-aware Open-vocabulary Vision Transformers (RO-ViT) - a contrastive image-text pretraining recipe to bridge the gap between image-level pretraining and open-vocabulary object detection. At the pretraining phase, we propose to randomly crop and resize regions of positional embeddings instead of using the whole image positional embeddings. This better matches the use of positional embeddings at region-level in the detection finetuning phase. In addition, we replace the common softmax cross entropy loss in contrastive learning with focal loss to better learn the informative yet difficult examples. Finally, we leverage recent advances in novel object proposals to improve open-vocabulary detection finetuning. We evaluate our full model on the LVIS and COCO open-vocabulary detection benchmarks and zero-shot transfer. RO-ViT achieves a state-of-the-art 32.1 AP_r on LVIS, surpassing the best existing approach by +5.8 points in addition to competitive zero-shot transfer detection. Surprisingly, RO-ViT improves the image-level representation as well and achieves the state of the art on 9 out of 12 metrics on COCO and Flickr image-text retrieval benchmarks, outperforming competitive approaches with larger models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers | TensorX