TensorX
返回文献探索

Paper · arXiv 2407.11691

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, Dahua Lin, Kai Chen

17 upvotesJuly 16, 2024arXiv 预印本
AI 摘要

VLMEvalKit is a PyTorch-based toolkit for evaluating large multi-modality models with a focus on vision-language models, offering a comprehensive framework for reproducible results and a leaderboard for research tracking.

multi-modality modelsPyTorchmulti-modal benchmarksdistributed inferenceprediction post-processingmetric calculationOpenVLM Leaderboard

Abstract

We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friendly and comprehensive framework for researchers and developers to evaluate existing multi-modality models and publish reproducible evaluation results. In VLMEvalKit, we implement over 70 different large multi-modality models, including both proprietary APIs and open-source models, as well as more than 20 different multi-modal benchmarks. By implementing a single interface, new models can be easily added to the toolkit, while the toolkit automatically handles the remaining workloads, including data preparation, distributed inference, prediction post-processing, and metric calculation. Although the toolkit is currently mainly used for evaluating large vision-language models, its design is compatible with future updates that incorporate additional modalities, such as audio and video. Based on the evaluation results obtained with the toolkit, we host OpenVLM Leaderboard, a comprehensive leaderboard to track the progress of multi-modality learning research. The toolkit is released at https://github.com/open-compass/VLMEvalKit and is actively maintained.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models | TensorX