TensorX
返回文献探索

Paper · arXiv 2509.19803

VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models

Guochao Jiang, Wenfeng Feng, Guofeng Quan, Chuzhan Hao, Yuewei Zhang, Guohua Liu, Hao Wang

122 upvotesSeptember 24, 2025arXiv 预印本
AI 摘要

A curriculum reinforcement learning framework dynamically adjusts training sample difficulty based on reward variance, improving LLM performance on mathematical reasoning tasks.

policy-based reinforcement learningrollout-based reinforcement learningGRPODAPOGSPORLVRcurriculum reinforcement learningVCRLmathematical reasoning tasksreward variance

Abstract

Policy-based reinforcement learning currently plays an important role in improving LLMs on mathematical reasoning tasks. However, existing rollout-based reinforcement learning methods (GRPO, DAPO, GSPO, etc.) fail to explicitly consider LLMs' learning ability for samples of different difficulty levels, which is contrary to the human cognitive process of mathematical reasoning tasks from easy to difficult. Intuitively, we find that the variance of the rollout group's reward in RLVR partly reflects the difficulty of the current sample for LLMs. Samples that are too easy or too difficult have a lower variance, while samples with moderate difficulty have a higher variance. Based on this, we propose VCRL, a curriculum reinforcement learning framework that dynamically controls the difficulty of training samples based on the variance of group rewards. Experiments on five mathematical benchmarks and two models reveal the advantages of VCRL over the current LLM RL baselines.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号