TensorX
返回文献探索

Paper · arXiv 2608.05802

On-Policy Delta Distillation for Multilingual Math Reasoning

Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han

32 upvotesAugust 6, 2026arXiv 预印本
AI 摘要

On-policy delta distillation improves multilingual mathematical reasoning and reduces cross-language performance gaps, though multilingual data is needed to preserve target-language outputs.

On-Policy DistillationOn-Policy Delta Distillationreinforcement learningLLM post-trainingmultilingualmathematical reasoningprobability gapQwen3

Abstract

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japanese. OPD^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD^2 consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
On-Policy Delta Distillation for Multilingual Math Reasoning | TensorX