TensorX
返回文献探索

Paper · arXiv 2606.04036

Self-Distilled Policy Gradient

Yifeng Liu, Shiyuan Zhang, Yifan Zhang, Quanquan Gu

28 upvotesJune 2, 2026arXiv 预印本
AI 摘要

A self-distilled policy-gradient framework combines on-policy self-distillation with verifier advantages and KL regularization to improve reinforcement learning stability and performance.

self-distillationpolicy-gradientreverse Kullback-Leibler divergenceon-policy learningverifier advantagesKL regularizationreinforcement learning

Abstract

On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense supervision for sparse-reward reinforcement learning. Actually, it can be instantiated as an auxiliary full-vocabulary student-to-teacher reverse Kullback-Leibler divergence loss. We therefore propose SDPG, a self-distilled policy-gradient framework that combines group-relative verifier advantages with normalized standard deviation, exact full-vocabulary on-policy self-distillation, as well as reference-policy KL regularization. Empirically, SDPG improves stability and performance over RLVR and self-distillation baselines. The code is available at https://github.com/lauyikfung/SDPG.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Self-Distilled Policy Gradient | TensorX