TensorX
返回文献探索

Paper · arXiv 2410.01257

HelpSteer2-Preference: Complementing Ratings with Preferences

Zhilin Wang, Alexander Bukharin, Olivier Delalleau, Daniel Egert, Gerald Shen, Jiaqi Zeng, Oleksii Kuchaiev, Yi Dong

26 upvotesOctober 2, 2024arXiv 预印本
AI 摘要

A novel combined Bradley-Terry and Regression reward modeling approach outperforms other models on RewardBench and aligns models to follow instructions in RLHF.

Brunswick-TerryRegressionpreference annotationshuman-written justificationshead-to-head comparisonreward modelingRLHFRewardBenchLlama-3.1-70B-InstructNVIDIA/Llama-3.1-Nemotron-70B-Reward

Abstract

Reward models are critical for aligning models to follow instructions, and are typically trained following one of two popular paradigms: Bradley-Terry style or Regression style. However, there is a lack of evidence that either approach is better than the other, when adequately matched for data. This is primarily because these approaches require data collected in different (but incompatible) formats, meaning that adequately matched data is not available in existing public datasets. To tackle this problem, we release preference annotations (designed for Bradley-Terry training) to complement existing ratings (designed for Regression style training) in the HelpSteer2 dataset. To improve data interpretability, preference annotations are accompanied with human-written justifications. Using this data, we conduct the first head-to-head comparison of Bradley-Terry and Regression models when adequately matched for data. Based on insights derived from such a comparison, we propose a novel approach to combine Bradley-Terry and Regression reward modeling. A Llama-3.1-70B-Instruct model tuned with this approach scores 94.1 on RewardBench, emerging top of more than 140 reward models as of 1 Oct 2024. We also demonstrate the effectiveness of this reward model at aligning models to follow instructions in RLHF. We open-source this dataset (CC-BY-4.0 license) at https://huggingface.co/datasets/nvidia/HelpSteer2 and openly release the trained Reward Model at https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Reward

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
HelpSteer2-Preference: Complementing Ratings with Preferences | TensorX