TensorX
返回文献探索

Paper · arXiv 2407.11522

FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models

Pengxiang Li, Zhi Gao, Bofei Zhang, Tao Yuan, Yuwei Wu, Mehrtash Harandi, Yunde Jia, Song-Chun Zhu, Qing Li

9 upvotesJuly 16, 2024arXiv 预印本
AI 摘要

FIRE, a feedback-refinement dataset, enhances VLMs' performance through user feedback, leading to a FIRE-LLaVA model that outperforms its untrained counterparts by 50% on a comprehensive benchmark.

Vision language modelsfeedback-refinementdatasetmulti-turn conversationsGPT-4Vbenchmarkevaluation settingsfine-tuningLLaVAmodelfeedbackuser-agent interactions

Abstract

Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is generated by GPT-4V, and FIRE-1M is freely generated via models trained on FIRE-100K. Then, we build FIRE-Bench, a benchmark to comprehensively evaluate the feedback-refining capability of VLMs, which contains 11K feedback-refinement conversations as the test data, two evaluation settings, and a model to provide feedback for VLMs. We develop the FIRE-LLaVA model by fine-tuning LLaVA on FIRE-100K and FIRE-1M, which shows remarkable feedback-refining capability on FIRE-Bench and outperforms untrained VLMs by 50%, making more efficient user-agent interactions and underscoring the significance of the FIRE dataset.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models | TensorX