TensorX
返回文献探索

Paper · arXiv 2402.03766

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, Shuang Xu, Fei Wei, Yang Yang, Xiaofei Sun, Yiming Hu, Xinyang Lin, Bo Zhang, Chunhua Shen

15 upvotesFebruary 6, 2024arXiv 预印本
AI 摘要

MobileVLM V2 enhances vision language models through improved architecture, training, and dataset curation, achieving performance on par with or better than larger models.

vision language modelsarchitectural designtraining schemedataset curationVLM benchmarks

Abstract

We introduce MobileVLM V2, a family of significantly improved vision language models upon MobileVLM, which proves that a delicate orchestration of novel architectural design, an improved training scheme tailored for mobile VLMs, and rich high-quality dataset curation can substantially benefit VLMs' performance. Specifically, MobileVLM V2 1.7B achieves better or on-par performance on standard VLM benchmarks compared with much larger VLMs at the 3B scale. Notably, our 3B model outperforms a large variety of VLMs at the 7B+ scale. Our models will be released at https://github.com/Meituan-AutoML/MobileVLM .

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号