TensorX
返回文献探索

Paper · arXiv 2605.30789

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

Yiming Ren, Yiran Xu, Zicheng Lin, Chufan Shi, Yukang Chen, Dingdong Wang, Tianhe Wu, Junjie Wang, Yujiu Yang, Yu Qiao, Ruihang Chu

27 upvotesJune 2, 2026arXiv 预印本
AI 摘要

Small-to-Large Policy Optimization framework uses smaller models as natural explorers to enhance policy diversity and improve large language model training efficiency.

Group Relative Policy Optimizationrollout diversitytoken-level randomnesspolicy-level diversitypass@ktemporal correlationgradient estimationsmall-to-large policy optimizationprogressive annealingmathematical reasoning benchmarks

Abstract

We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomness, which may introduce step-wise noise and lead to incoherent trajectories. We uncover that smaller models within the same model family inherently exhibit higher policy-level diversity, indicated by their superior pass@k relative to larger counterparts as sample counts increase. Unlike token-level noise, this diversity is temporally correlated, preserves logical consistency, and provides structured exploration signals for gradient estimation. We thus propose S2L-PO (Small-to-Large Policy Optimization), a framework that leverages fixed small models as natural explorers to train larger models. To balance exploration and exploitation, we design a progressive annealing strategy that transitions from offline small-model rollouts to the large learner's own sampling. This shift elegantly avoids mid-training performance drops caused by the small model's capacity limits, achieving faster convergence and unlocking a higher performance ceiling. S2L-PO improves accuracy on diverse mathematical reasoning benchmarks (e.g., +8.8% on AIME 24 using a 1.7B explorer to guide the 8B model) while reducing rollout compute.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO | TensorX