TensorX
返回文献探索

Paper · arXiv 2401.10020

Self-Rewarding Language Models

Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, Jason Weston

156 upvotesJanuary 18, 2024arXiv 预印本
AI 摘要

A study on Self-Rewarding Language Models shows that using LLM-as-a-Judge prompting for iterative DPO training enhances both instruction-following and self-reward generation, leading to superior performance compared to existing systems.

Self-Rewarding Language ModelsLLM-as-a-JudgeIterative DPO trainingAlpacaEval 2.0Claude 2Gemini ProGPT-4 0613

Abstract

We posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these separate frozen reward models cannot then learn to improve during LLM training. In this work, we study Self-Rewarding Language Models, where the language model itself is used via LLM-as-a-Judge prompting to provide its own rewards during training. We show that during Iterative DPO training that not only does instruction following ability improve, but also the ability to provide high-quality rewards to itself. Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613. While only a preliminary study, this work opens the door to the possibility of models that can continually improve in both axes.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号