TensorX
返回文献探索

Paper · arXiv 2402.00742

Transforming and Combining Rewards for Aligning Large Language Models

Zihao Wang, Chirag Nagpal, Jonathan Berant, Jacob Eisenstein, Alex D'Amour, Sanmi Koyejo, Victor Veitch

12 upvotesFebruary 1, 2024arXiv 预印本
AI 摘要

The study examines the transformation and aggregation of reward models to align language models with human preferences, identifying a method that improves poor outputs and aggregates rewards logically.

reward modelpreference datapreference rankingalignment procedureprobabilistic interpretationBradley-Terry preference modelsreward transformationpoorly-performing outputsreward hackingunderfittingaggregated rewardslogical conjunctionRLHFhelpfulharmless

Abstract

A common approach for aligning language models to human preferences is to first learn a reward model from preference data, and then use this reward model to update the language model. We study two closely related problems that arise in this approach. First, any monotone transformation of the reward model preserves preference ranking; is there a choice that is ``better'' than others? Second, we often wish to align language models to multiple properties: how should we combine multiple reward models? Using a probabilistic interpretation of the alignment procedure, we identify a natural choice for transformation for (the common case of) rewards learned from Bradley-Terry preference models. This derived transformation has two important properties. First, it emphasizes improving poorly-performing outputs, rather than outputs that already score well. This mitigates both underfitting (where some prompts are not improved) and reward hacking (where the model learns to exploit misspecification of the reward model). Second, it enables principled aggregation of rewards by linking summation to logical conjunction: the sum of transformed rewards corresponds to the probability that the output is ``good'' in all measured properties, in a sense we make precise. Experiments aligning language models to be both helpful and harmless using RLHF show substantial improvements over the baseline (non-transformed) approach.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Transforming and Combining Rewards for Aligning Large Language Models | TensorX