TensorX
返回文献探索

Paper · arXiv 2607.04751

Trust Region Policy Distillation

Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang

37 upvotesJuly 6, 2026arXiv 预印本
AI 摘要

Trust Region Policy Distillation stabilizes on-policy distillation via a proximal teacher, reducing gradient variance and improving convergence for mathematical reasoning tasks without extra computation.

Trust Region Policy DistillationOn-Policy Distillationproximal teachergradient varianceglobal convergence analysismonotonic improvement bound

Abstract

Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Trust Region Policy Distillation | TensorX