TensorX
返回文献探索

Paper · arXiv 2310.03716

A Long Way to Go: Investigating Length Correlations in RLHF

Prasann Singhal, Tanya Goyal, Jiacheng Xu, Greg Durrett

10 upvotesOctober 5, 2023arXiv 预印本
AI 摘要

Optimizing for response length significantly contributes to the improvements observed in Reinforcement Learning from Human Feedback (RLHF) when aligning large language models for helpfulness tasks.

Reinforcement Learning from Human Feedback (RLHF)reward modelspreference datasetshelpfulnessmulti-turn dialogueweb question answeringsummarizationresponse lengthdownstream improvementsrl

Abstract

Great successes have been reported using Reinforcement Learning from Human Feedback (RLHF) to align large language models. Open-source preference datasets and reward models have enabled wider experimentation beyond generic chat settings, particularly to make systems more "helpful" for tasks like web question answering, summarization, and multi-turn dialogue. When optimizing for helpfulness, RLHF has been consistently observed to drive models to produce longer outputs. This paper demonstrates that optimizing for response length is a significant factor behind RLHF's reported improvements in these settings. First, we study the relationship between reward and length for reward models trained on three open-source preference datasets for helpfulness. Here, length correlates strongly with reward, and improvements in reward score are driven in large part by shifting the distribution over output lengths. We then explore interventions during both RL and reward model learning to see if we can achieve the same downstream improvements as RLHF without increasing length. While our interventions mitigate length increases, they aren't uniformly effective across settings. Furthermore, we find that even running RLHF with a reward based solely on length can reproduce most of the downstream improvements over the initial policy model, showing that reward models in these settings have a long way to go.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
A Long Way to Go: Investigating Length Correlations in RLHF | TensorX