TensorX
返回文献探索

Paper · arXiv 2309.00267

RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, Abhinav Rastogi

53 upvotesSeptember 1, 2023arXiv 预印本
AI 摘要

Reinforcement learning from AI feedback (RLAIF) achieves similar performance to reinforcement learning from human feedback (RLHF) in aligning large language models with human preferences, offering a scalable alternative.

Reinforcement learning from human feedbackRLHFReinforcement learning from AI FeedbackRLAIFlarge language modelsLLMshuman preferencespreference labelshead-to-head comparisonsummarizationhuman evaluatorsbaseline supervised fine-tuned modelhuman-level performance

Abstract

Reinforcement learning from human feedback (RLHF) is effective at aligning large language models (LLMs) to human preferences, but gathering high quality human preference labels is a key bottleneck. We conduct a head-to-head comparison of RLHF vs. RL from AI Feedback (RLAIF) - a technique where preferences are labeled by an off-the-shelf LLM in lieu of humans, and we find that they result in similar improvements. On the task of summarization, human evaluators prefer generations from both RLAIF and RLHF over a baseline supervised fine-tuned model in ~70% of cases. Furthermore, when asked to rate RLAIF vs. RLHF summaries, humans prefer both at equal rates. These results suggest that RLAIF can yield human-level performance, offering a potential solution to the scalability limitations of RLHF.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号