TensorX
返回文献探索

Paper · arXiv 2410.18451

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Chris Yuhao Liu, Liang Zeng, Jiacai Liu, Rui Yan, Jujie He, Chaojie Wang, Shuicheng Yan, Yang Liu, Yahui Zhou

20 upvotesOctober 24, 2024arXiv 预印本
AI 摘要

Curated preference datasets and selection strategies enhanced reward models, leading to improved performance on the RewardBench leaderboard.

Kywork-Rewarddata selectionfiltering strategiesSkywork-Reward-Gemma-27BSkywork-Reward-Llama-3.1-8BRewardBench

Abstract

In this report, we introduce a collection of methods to enhance reward modeling for LLMs, focusing specifically on data-centric techniques. We propose effective data selection and filtering strategies for curating high-quality open-source preference datasets, culminating in the Skywork-Reward data collection, which contains only 80K preference pairs -- significantly smaller than existing datasets. Using this curated dataset, we developed the Skywork-Reward model series -- Skywork-Reward-Gemma-27B and Skywork-Reward-Llama-3.1-8B -- with the former currently holding the top position on the RewardBench leaderboard. Notably, our techniques and datasets have directly enhanced the performance of many top-ranked models on RewardBench, highlighting the practical impact of our contributions in real-world preference learning applications.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs | TensorX