TensorX
返回文献探索

Paper · arXiv 2409.11564

Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey

Genta Indra Winata, Hanyang Zhao, Anirban Das, Wenpin Tang, David D. Yao, Shi-Xiong Zhang, Sambit Sahu

20 upvotesSeptember 17, 2024arXiv 预印本
AI 摘要

Preference tuning in deep generative models integrates human feedback across modalities using reinforcement learning to align model outputs with human preferences, with a focus on methodologies and future directions.

preference tuningdeep generative modelshuman preferencesreinforcement learningpreference tuning taskspolicy approacheslanguagespeechvisionevaluation methodsmodel alignment

Abstract

Preference tuning is a crucial process for aligning deep generative models with human preferences. This survey offers a thorough overview of recent advancements in preference tuning and the integration of human feedback. The paper is organized into three main sections: 1) introduction and preliminaries: an introduction to reinforcement learning frameworks, preference tuning tasks, models, and datasets across various modalities: language, speech, and vision, as well as different policy approaches, 2) in-depth examination of each preference tuning approach: a detailed analysis of the methods used in preference tuning, and 3) applications, discussion, and future directions: an exploration of the applications of preference tuning in downstream tasks, including evaluation methods for different modalities, and an outlook on future research directions. Our objective is to present the latest methodologies in preference tuning and model alignment, enhancing the understanding of this field for researchers and practitioners. We hope to encourage further engagement and innovation in this area.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号