TensorX
返回文献探索

Paper · arXiv 2411.10669

Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts

Jinqiang Long, Yanqi Dai, Guoxing Yang, Hongpeng Lin, Nanyi Fei, Yizhao Gao, Zhiwu Lu

10 upvotesNovember 16, 2024arXiv 预印本
AI 摘要

Awaker2.5-VL uses a Mixture of Experts architecture with low-rank adaptation to improve performance in various multimodal tasks.

Multimodal Large Language ModelsMLLMsVQADetectionOCRChartQAmulti-task conflictMixture of ExpertsMoElow-rank adaptationLoRA

Abstract

As the research of Multimodal Large Language Models (MLLMs) becomes popular, an advancing MLLM model is typically required to handle various textual and visual tasks (e.g., VQA, Detection, OCR, and ChartQA) simultaneously for real-world applications. However, due to the significant differences in representation and distribution among data from various tasks, simply mixing data of all tasks together leads to the well-known``multi-task conflict" issue, resulting in performance degradation across various tasks. To address this issue, we propose Awaker2.5-VL, a Mixture of Experts~(MoE) architecture suitable for MLLM, which acquires the multi-task capabilities through multiple sparsely activated experts. To speed up the training and inference of Awaker2.5-VL, each expert in our model is devised as a low-rank adaptation (LoRA) structure. Extensive experiments on multiple latest benchmarks demonstrate the effectiveness of Awaker2.5-VL. The code and model weight are released in our Project Page: https://github.com/MetabrainAGI/Awaker.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts | TensorX