TensorX
返回文献探索

Paper · arXiv 2410.02712

LLaVA-Critic: Learning to Evaluate Multimodal Models

Tianyi Xiong, Xiyao Wang, Dong Guo, Qinghao Ye, Haoqi Fan, Quanquan Gu, Heng Huang, Chunyuan Li

37 upvotesOctober 3, 2024arXiv 预印本
AI 摘要

LLaVA-Critic, an open-source large multimodal model, effectively evaluates multimodal tasks and provides reliable scores, surpassing GPT models, and enhances preference learning for model alignment.

large multimodal modelLMMinstruction-following datasetevaluation criteriaLMM-as-a-JudgeGPT modelsevaluation benchmarksPreference Learningreward signalsmodel alignmentsuperhuman alignment feedback mechanisms

Abstract

We introduce LLaVA-Critic, the first open-source large multimodal model (LMM) designed as a generalist evaluator to assess performance across a wide range of multimodal tasks. LLaVA-Critic is trained using a high-quality critic instruction-following dataset that incorporates diverse evaluation criteria and scenarios. Our experiments demonstrate the model's effectiveness in two key areas: (1) LMM-as-a-Judge, where LLaVA-Critic provides reliable evaluation scores, performing on par with or surpassing GPT models on multiple evaluation benchmarks; and (2) Preference Learning, where it generates reward signals for preference learning, enhancing model alignment capabilities. This work underscores the potential of open-source LMMs in self-critique and evaluation, setting the stage for future research into scalable, superhuman alignment feedback mechanisms for LMMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLaVA-Critic: Learning to Evaluate Multimodal Models | TensorX