TensorX
返回文献探索

Paper · arXiv 2505.07787

Learning from Peers in Reasoning Models

Tongxu Luo, Wenyu Du, Jiaxi Bi, Stephen Chung, Zhengyang Tang, Hao Yang, Min Zhang, Benyou Wang

45 upvotesMay 12, 2025arXiv 预印本
AI 摘要

LeaP, a peer interaction mechanism for large reasoning models, improves error correction and performance across various math benchmarks by enabling collaborative reasoning paths.

Learning from Peers (LeaP)Prefix Dominance Trappeer interactionrouting mechanismintermediate reasoningfine-tuningLeaP-Terror correctionQwQ-32BDeepSeek-R1-671BDeepSeek-R1-Distill-Qwen-14BAIME 2024AIME 2025AIMO 2025GPQA Diamond

Abstract

Large Reasoning Models (LRMs) have the ability to self-correct even when they make mistakes in their reasoning paths. However, our study reveals that when the reasoning process starts with a short but poor beginning, it becomes difficult for the model to recover. We refer to this phenomenon as the "Prefix Dominance Trap". Inspired by psychological findings that peer interaction can promote self-correction without negatively impacting already accurate individuals, we propose **Learning from Peers** (LeaP) to address this phenomenon. Specifically, every tokens, each reasoning path summarizes its intermediate reasoning and shares it with others through a routing mechanism, enabling paths to incorporate peer insights during inference. However, we observe that smaller models sometimes fail to follow summarization and reflection instructions effectively. To address this, we fine-tune them into our **LeaP-T** model series. Experiments on AIME 2024, AIME 2025, AIMO 2025, and GPQA Diamond show that LeaP provides substantial improvements. For instance, QwQ-32B with LeaP achieves nearly 5 absolute points higher than the baseline on average, and surpasses DeepSeek-R1-671B on three math benchmarks with an average gain of 3.3 points. Notably, our fine-tuned LeaP-T-7B matches the performance of DeepSeek-R1-Distill-Qwen-14B on AIME 2024. In-depth analysis reveals LeaP's robust error correction by timely peer insights, showing strong error tolerance and handling varied task difficulty. LeaP marks a milestone by enabling LRMs to collaborate during reasoning. Our code, datasets, and models are available at https://learning-from-peers.github.io/ .

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号