TensorX
返回文献探索

Paper · arXiv 2510.26788

Defeating the Training-Inference Mismatch via FP16

Penghui Qi, Zichen Liu, Xiangxin Zhou, Tianyu Pang, Chao Du, Wee Sun Lee, Min Lin

32 upvotesOctober 30, 2025arXiv 预印本
AI 摘要

Using FP16 precision in reinforcement learning fine-tuning of large language models improves stability, convergence, and performance by addressing numerical mismatches.

reinforcement learninglarge language modelsnumerical mismatchtraining policiesinference policiesfloating point precisionBF16FP16stabilityconvergenceperformance

Abstract

Reinforcement learning (RL) fine-tuning of large language models (LLMs) often suffers from instability due to the numerical mismatch between the training and inference policies. While prior work has attempted to mitigate this issue through algorithmic corrections or engineering alignments, we show that its root cause lies in the floating point precision itself. The widely adopted BF16, despite its large dynamic range, introduces large rounding errors that breaks the consistency between training and inference. In this work, we demonstrate that simply reverting to FP16 effectively eliminates this mismatch. The change is simple, fully supported by modern frameworks with only a few lines of code change, and requires no modification to the model architecture or learning algorithm. Our results suggest that using FP16 uniformly yields more stable optimization, faster convergence, and stronger performance across diverse tasks, algorithms and frameworks. We hope these findings motivate a broader reconsideration of precision trade-offs in RL fine-tuning.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Defeating the Training-Inference Mismatch via FP16 | TensorX