TensorX
返回文献探索

Paper · arXiv 2506.08007

Reinforcement Pre-Training

Qingxiu Dong, Li Dong, Yao Tang, Tianzhu Ye, Yutao Sun, Zhifang Sui, Furu Wei

265 upvotesJune 9, 2025arXiv 预印本
AI 摘要

Reinforcement Pre-Training (RPT) improves language model accuracy through reinforcement learning and offers a scalable method for leveraging text data for general-purpose RL.

Reinforcement Pre-Training (RPT)next-token predictionreasoning taskreinforcement learning (RL)verifiable rewardslanguage modeling accuracyreinforcement fine-tuningscaling curves

Abstract

In this work, we introduce Reinforcement Pre-Training (RPT) as a new scaling paradigm for large language models and reinforcement learning (RL). Specifically, we reframe next-token prediction as a reasoning task trained using RL, where it receives verifiable rewards for correctly predicting the next token for a given context. RPT offers a scalable method to leverage vast amounts of text data for general-purpose RL, rather than relying on domain-specific annotated answers. By incentivizing the capability of next-token reasoning, RPT significantly improves the language modeling accuracy of predicting the next tokens. Moreover, RPT provides a strong pre-trained foundation for further reinforcement fine-tuning. The scaling curves show that increased training compute consistently improves the next-token prediction accuracy. The results position RPT as an effective and promising scaling paradigm to advance language model pre-training.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Reinforcement Pre-Training | TensorX