TensorX
返回文献探索

Paper · arXiv 2503.22230

Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

Wei Shen, Guanlin Liu, Zheng Wu, Ruofei Zhu, Qingping Yang, Chao Xin, Yu Yue, Lin Yan

45 upvotesMarch 28, 2025arXiv 预印本
AI 摘要

This study enhances RLHF performance by addressing reward hacking and response diversity through a hybrid reward system and a novel prompt-selection method.

Reinforcement Learning from Human Feedback (RLHF)reasoning task verifiers (RTV)generative reward model (GenRM)Pre-PPOreward hackingresponse diversitymathematical taskscoding tasksSFT Best-of-N responses

Abstract

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning large language models with human preferences. While recent research has focused on algorithmic improvements, the importance of prompt-data construction has been overlooked. This paper addresses this gap by exploring data-driven bottlenecks in RLHF performance scaling, particularly reward hacking and decreasing response diversity. We introduce a hybrid reward system combining reasoning task verifiers (RTV) and a generative reward model (GenRM) to mitigate reward hacking. We also propose a novel prompt-selection method, Pre-PPO, to maintain response diversity and enhance learning effectiveness. Additionally, we find that prioritizing mathematical and coding tasks early in RLHF training significantly improves performance. Experiments across two model sizes validate our methods' effectiveness and scalability. Results show that RTV is most resistant to reward hacking, followed by GenRM with ground truth, and then GenRM with SFT Best-of-N responses. Our strategies enable rapid capture of subtle task-specific distinctions, leading to substantial improvements in overall RLHF performance. This work highlights the importance of careful data construction and provides practical methods to overcome performance barriers in RLHF.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号