TensorX
返回文献探索

Paper · arXiv 2502.14768

Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Tian Xie, Zitian Gao, Qingnan Ren, Haoming Luo, Yuqian Hong, Bryan Dai, Joey Zhou, Kai Qiu, Zhirong Wu, Chong Luo

47 upvotesFebruary 20, 2025arXiv 预印本
AI 摘要

A rule-based reinforcement learning system for large reasoning models uses synthetic logic puzzles to develop advanced reasoning skills and demonstrates generalization on complex math benchmarks.

rule-based reinforcement learningRLsynthetic logic puzzlesformat reward functionstable convergencereflectionverificationsummarizationmath benchmarksAIMEAMC

Abstract

Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in large reasoning models. To analyze reasoning dynamics, we use synthetic logic puzzles as training data due to their controllable complexity and straightforward answer verification. We make some key technical contributions that lead to effective and stable RL training: a system prompt that emphasizes the thinking and answering process, a stringent format reward function that penalizes outputs for taking shortcuts, and a straightforward training recipe that achieves stable convergence. Our 7B model develops advanced reasoning skills-such as reflection, verification, and summarization-that are absent from the logic corpus. Remarkably, after training on just 5K logic problems, it demonstrates generalization abilities to the challenging math benchmarks AIME and AMC.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号