TensorX
返回文献探索

Paper · arXiv 2509.18849

MAPO: Mixed Advantage Policy Optimization

Wenke Huang, Quan Zhang, Yiyang Fang, Jian Liang, Xuankun Rong, Huanjin Yao, Guancheng Wan, Ke Liang, Wenwen He, Mingjun Li, Leszek Rutkowski, Mang Ye, Bo Du, Dacheng Tao

27 upvotesSeptember 23, 2025arXiv 预印本
AI 摘要

Mixed Advantage Policy Optimization (MAPO) dynamically reweights the advantage function to improve trajectory ranking in reinforcement learning for foundation models.

Group Relative Policy Optimization (GRPO)advantage functiontrajectory importanceadvantage reversionadvantage mirror problemsMixed Advantage Policy Optimization (MAPO)trajectory certaintyadvantage percent deviationsample-specific characteristics

Abstract

Recent advances in reinforcement learning for foundation models, such as Group Relative Policy Optimization (GRPO), have significantly improved the performance of foundation models on reasoning tasks. Notably, the advantage function serves as a central mechanism in GRPO for ranking the trajectory importance. However, existing explorations encounter both advantage reversion and advantage mirror problems, which hinder the reasonable advantage allocation across different query samples. In this work, we propose an easy but effective GRPO strategy, Mixed Advantage Policy Optimization (MAPO). We reveal that the trajectory appears with different certainty and propose the advantage percent deviation for samples with high-certainty trajectories. Furthermore, we dynamically reweight the advantage function for samples with varying trajectory certainty, thereby adaptively configuring the advantage function to account for sample-specific characteristics. Comparison with related state-of-the-art methods, along with ablation studies on different advantage variants, validates the effectiveness of our approach.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MAPO: Mixed Advantage Policy Optimization | TensorX