TensorX
返回文献探索

Paper · arXiv 2408.16357

Law of Vision Representation in MLLMs

Shijia Yang, Bohan Zhai, Quanzeng You, Jianbo Yuan, Hongxia Yang, Chenfeng Xu

95 upvotesAugust 29, 2024arXiv 预印本
AI 摘要

Correlation between cross-modal alignment and vision representation improves performance in multimodal large language models, enabling identification and training of optimal vision representation with reduced computational cost.

cross-modal alignmentvision representationMLLMsAC scoremultimodal large language models

Abstract

We present the "Law of Vision Representation" in multimodal large language models (MLLMs). It reveals a strong correlation between the combination of cross-modal alignment, correspondence in vision representation, and MLLM performance. We quantify the two factors using the cross-modal Alignment and Correspondence score (AC score). Through extensive experiments involving thirteen different vision representation settings and evaluations across eight benchmarks, we find that the AC score is linearly correlated to model performance. By leveraging this relationship, we are able to identify and train the optimal vision representation only, which does not require finetuning the language model every time, resulting in a 99.7% reduction in computational cost.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号