TensorX
返回文献探索

Paper · arXiv 2305.10425

SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, Peter J. Liu

7 upvotesMay 17, 2023arXiv 预印本
AI 摘要

Sequence Likelihood Calibration (SLiC) is shown to be an effective and simpler alternative to Reinforcement Learning from Human Feedback (RLHF) for learning from human preferences in language models.

Reinforcement Learning from Human Feedback (RLHF)Sequence Likelihood Calibration (SLiC)human feedbackhuman preferencesPPOsupervised fine-tuning

Abstract

Learning from human feedback has been shown to be effective at aligning language models with human preferences. Past work has often relied on Reinforcement Learning from Human Feedback (RLHF), which optimizes the language model using reward scores assigned from a reward model trained on human preference data. In this work we show how the recently introduced Sequence Likelihood Calibration (SLiC), can also be used to effectively learn from human preferences (SLiC-HF). Furthermore, we demonstrate this can be done with human feedback data collected for a different model, similar to off-policy, offline RL data. Automatic and human evaluation experiments on the TL;DR summarization task show that SLiC-HF significantly improves supervised fine-tuning baselines. Furthermore, SLiC-HF presents a competitive alternative to the PPO RLHF implementation used in past work while being much simpler to implement, easier to tune and more computationally efficient in practice.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SLiC-HF: Sequence Likelihood Calibration with Human Feedback | TensorX