TensorX
返回文献探索

Paper · arXiv 2503.03746

Process-based Self-Rewarding Language Models

Shimao Zhang, Xiao Liu, Xin Zhang, Junxiao Liu, Zheheng Luo, Shujian Huang, Yeyun Gong

39 upvotesMarch 5, 2025arXiv 预印本
AI 摘要

A Process-based Self-Rewarding pipeline enhances LLMs' mathematical reasoning by iteratively evaluating and optimizing their own outputs.

Large Language Modelsself-rewardingLLM-as-a-Judgestep-wise preference optimizationmathematical reasoningProcess-based Self-Rewarding

Abstract

Large Language Models have demonstrated outstanding performance across various downstream tasks and have been widely applied in multiple scenarios. Human-annotated preference data is used for training to further improve LLMs' performance, which is constrained by the upper limit of human performance. Therefore, Self-Rewarding method has been proposed, where LLMs generate training data by rewarding their own outputs. However, the existing self-rewarding paradigm is not effective in mathematical reasoning scenarios and may even lead to a decline in performance. In this work, we propose the Process-based Self-Rewarding pipeline for language models, which introduces long-thought reasoning, step-wise LLM-as-a-Judge, and step-wise preference optimization within the self-rewarding paradigm. Our new paradigm successfully enhances the performance of LLMs on multiple mathematical reasoning benchmarks through iterative Process-based Self-Rewarding, demonstrating the immense potential of self-rewarding to achieve LLM reasoning that may surpass human capabilities.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号