TensorX
返回文献探索

Paper · arXiv 2503.18931

CoMP: Continual Multimodal Pre-training for Vision Foundation Models

Yitong Chen, Lingchen Meng, Wujian Peng, Zuxuan Wu, Yu-Gang Jiang

31 upvotesMarch 24, 2025arXiv 预印本
AI 摘要

CoMP, a multimodal pre-training pipeline with Continual Rotary Position Embedding and Alignment Loss, enhances visual representations for various tasks by aligning with language models.

Pre-trained Vision Foundation ModelsContinual Rotary Position EmbeddingMultimodal pre-training pipelineAlignment Losslanguage prototypesmultimodal understandingChartQADocVQAImageNet-1KADE20Kfrozen chunk evaluation

Abstract

Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this paper, we continually pre-train prevailing VFMs in a multimodal manner such that they can effortlessly process visual inputs of varying sizes and produce visual representations that are more aligned with language representations, regardless of their original pre-training process. To this end, we introduce CoMP, a carefully designed multimodal pre-training pipeline. CoMP uses a Continual Rotary Position Embedding to support native resolution continual pre-training, and an Alignment Loss between visual and textual features through language prototypes to align multimodal representations. By three-stage training, our VFMs achieve remarkable improvements not only in multimodal understanding but also in other downstream tasks such as classification and segmentation. Remarkably, CoMP-SigLIP achieves scores of 66.7 on ChartQA and 75.9 on DocVQA with a 0.5B LLM, while maintaining an 87.4% accuracy on ImageNet-1K and a 49.5 mIoU on ADE20K under frozen chunk evaluation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
CoMP: Continual Multimodal Pre-training for Vision Foundation Models | TensorX