TensorX
返回文献探索

Paper · arXiv 2501.04682

Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Though

Violet Xiang, Charlie Snell, Kanishk Gandhi, Alon Albalak, Anikait Singh, Chase Blagden, Duy Phung, Rafael Rafailov, Nathan Lile, Dakota Mahan, Louis Castricato, Jan-Philipp Franken, Nick Haber, Chelsea Finn

99 upvotesJanuary 8, 2025arXiv 预印本
AI 摘要

The Meta Chain-of-Thought framework extends traditional CoT by explicitly modeling reasoning, enabling more powerful and human-like reasoning in LLMs through process supervision, synthetic data generation, and reinforcement learning.

Meta Chain-of-ThoughtMeta-CoTChain-of-ThoughtCoTin-context searchprocess supervisionsynthetic data generationsearch algorithmsinstruction tuninglinearized search tracesreinforcement learningscaling lawsverifier rolesreasoning algorithmsLLMs

Abstract

We propose a novel framework, Meta Chain-of-Thought (Meta-CoT), which extends traditional Chain-of-Thought (CoT) by explicitly modeling the underlying reasoning required to arrive at a particular CoT. We present empirical evidence from state-of-the-art models exhibiting behaviors consistent with in-context search, and explore methods for producing Meta-CoT via process supervision, synthetic data generation, and search algorithms. Finally, we outline a concrete pipeline for training a model to produce Meta-CoTs, incorporating instruction tuning with linearized search traces and reinforcement learning post-training. Finally, we discuss open research questions, including scaling laws, verifier roles, and the potential for discovering novel reasoning algorithms. This work provides a theoretical and practical roadmap to enable Meta-CoT in LLMs, paving the way for more powerful and human-like reasoning in artificial intelligence.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号