TensorX
返回文献探索

Paper · arXiv 2404.01258

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander Hauptmann, Yonatan Bisk, Yiming Yang

11 upvotesApril 1, 2024arXiv 预印本
AI 摘要

A novel framework using video captions as a proxy enhances preference modeling and performance of video LMMs by aligning with GPT-4V's reward mechanism through direct preference optimization.

direct preference optimization (DPO)large language model (LLM)video instruction-followingreward modelslarge multimodal models (LMMs)video captionsvideo Question Answering (QA)GPT-4V

Abstract

Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However, in tasks involving video instruction-following, providing informative feedback, especially for detecting hallucinations in generated responses, remains a significant challenge. Previous studies have explored using large large multimodal models (LMMs) as reward models to guide preference modeling, but their ability to accurately assess the factuality of generated responses compared to corresponding videos has not been conclusively established. This paper introduces a novel framework that utilizes detailed video captions as a proxy of video content, enabling language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions. Our approach demonstrates robust alignment with OpenAI GPT-4V model's reward mechanism, which directly takes video frames as input. Furthermore, we show that applying this tailored reward through DPO significantly improves the performance of video LMMs on video QA tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward | TensorX