TensorX
返回文献探索

Paper · arXiv 2406.11839

mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Fei Wang, Wenxuan Zhou, James Y. Huang, Nan Xu, Sheng Zhang, Hoifung Poon, Muhao Chen

40 upvotesJune 17, 2024arXiv 预印本
AI 摘要

mDPO, a multimodal preference optimization method, addresses the unconditional preference problem and enhances model performance in multimodal large language models.

direct preference optimization (DPO)large language model (LLM)multimodal preference optimizationunconditional preference problemmultimodal DPOimage conditionreward anchorrelative preference optimizationhallucination

Abstract

Direct preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment. Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement. Through a comparative experiment, we identify the unconditional preference problem in multimodal preference optimization, where the model overlooks the image condition. To address this problem, we propose mDPO, a multimodal DPO objective that prevents the over-prioritization of language-only preferences by also optimizing image preference. Moreover, we introduce a reward anchor that forces the reward to be positive for chosen responses, thereby avoiding the decrease in their likelihood -- an intrinsic problem of relative preference optimization. Experiments on two multimodal LLMs of different sizes and three widely used benchmarks demonstrate that mDPO effectively addresses the unconditional preference problem in multimodal preference optimization and significantly improves model performance, particularly in reducing hallucination.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
mDPO: Conditional Preference Optimization for Multimodal Large Language Models | TensorX