TensorX
返回文献探索

Paper · arXiv 2609.07108

Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training

Zili Wang, Zhaopeng Qiu, Yuekai Zhang, Shuang Yu, Junjie Lai

28 upvotesSeptember 7, 2026arXiv 预印本
AI 摘要

A system for large-scale online draft co-training accelerates speculative decoding in RL post-training by extending context-parallel attention and adding cross-stage feature transport.

speculative decodingreinforcement learning post-trainingonline co-trainingbranch attentioncontext-parallelzigzag ring attentionpipeline-parallelTapChannelrollout generation

Abstract

Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft's accuracy, yielding greater speedups. However, scaling this approach to co-training on large models with long contexts poses two obstacles: (1) branch attention is unsupported by standard causal context-parallel (CP) implementations, and (2) target features span across pipeline-parallel (PP) stages. We address both with an end-to-end system for large-scale online draft co-training. For CP, we extend packed, load-balanced zigzag ring attention by merging rank-local branch attention with causal main-sequence attention. For PP, TapChannel transports intermediate target features across stages via a separate path, leaving the pipeline schedule unaffected. Experiments demonstrate that co-trained drafts closely track the policy baseline while delivering substantial rollout and end-to-end speedups across model scales up to 122B. Our CP design achieves strong scaling at 256K tokens with significant memory savings over prior work, and our PP transport incurs modest overhead. Code can be found at https://github.com/NVIDIA-NeMo/RL/issues/3698.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号