TensorX
返回文献探索

Paper · arXiv 2504.04022

Rethinking Reflection in Pre-Training

Essential AI, Darsh J Shah, Peter Rushton, Somanshu Singla, Mohit Parmar, Kurt Smith, Yash Vanjani, Ashish Vaswani, Adarsh Chaluvaraju, Andrew Hojel, Andrew Ma, Anil Thomas, Anthony Polloreno, Ashish Tanwer, Burhan Drak Sibai, Divya S Mansingka, Divya Shivaprasad, Ishaan Shah, Karl Stratos, Khoi Nguyen, Michael Callahan, Michael Pust, Mrinal Iyer, Philip Monk, Platon Mazarakis, Ritvik Kapila, Saurabh Srivastava, Tim Romanski

80 upvotesApril 5, 2025arXiv 预印本
AI 摘要

Models exhibit self-correcting ability during pre-training by recognizing and addressing errors in their reasoning, a skill that improves over time.

language modelself-reflectionreinforcement learningpre-trainingchains-of-thoughtself-correcting abilityOLMo2-7Btokensself-reflection tasks

Abstract

A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develops during reinforcement learning, we show that it actually begins to emerge much earlier - during the model's pre-training. To study this, we introduce deliberate errors into chains-of-thought and test whether the model can still arrive at the correct answer by recognizing and correcting these mistakes. By tracking performance across different stages of pre-training, we observe that this self-correcting ability appears early and improves steadily over time. For instance, an OLMo2-7B model pre-trained on 4 trillion tokens displays self-correction on our six self-reflection tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号