TensorX
返回文献探索

Paper · arXiv 2405.19332

Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Shenao Zhang, Donghan Yu, Hiteshi Sharma, Ziyi Yang, Shuohang Wang, Hany Hassan, Zhaoran Wang

22 upvotesMay 29, 2024arXiv 预印本
AI 摘要

A bilevel objective for self-exploration in preference optimization enhances Large Language Models' performance by systematically exploring diverse responses and reducing bias towards unseen extrapolations.

Reinforcement Learning from Human FeedbackLarge Language Modelsreward modelsonline feedback collectionbilevel objectivereparameterized reward functionSelf-Exploring Language ModelsDirect Preference Optimizationinstruction-following benchmarksacademic benchmarks

Abstract

Preference optimization, particularly through Reinforcement Learning from Human Feedback (RLHF), has achieved significant success in aligning Large Language Models (LLMs) to adhere to human intentions. Unlike offline alignment with a fixed dataset, online feedback collection from humans or AI on model generations typically leads to more capable reward models and better-aligned LLMs through an iterative process. However, achieving a globally accurate reward model requires systematic exploration to generate diverse responses that span the vast space of natural language. Random sampling from standard reward-maximizing LLMs alone is insufficient to fulfill this requirement. To address this issue, we propose a bilevel objective optimistically biased towards potentially high-reward responses to actively explore out-of-distribution regions. By solving the inner-level problem with the reparameterized reward function, the resulting algorithm, named Self-Exploring Language Models (SELM), eliminates the need for a separate RM and iteratively updates the LLM with a straightforward objective. Compared to Direct Preference Optimization (DPO), the SELM objective reduces indiscriminate favor of unseen extrapolations and enhances exploration efficiency. Our experimental results demonstrate that when finetuned on Zephyr-7B-SFT and Llama-3-8B-Instruct models, SELM significantly boosts the performance on instruction-following benchmarks such as MT-Bench and AlpacaEval 2.0, as well as various standard academic benchmarks in different settings. Our code and models are available at https://github.com/shenao-zhang/SELM.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment | TensorX