TensorX
返回文献探索

Paper · arXiv 2504.10903

Efficient Reasoning Models: A Survey

Sicheng Feng, Gongfan Fang, Xinyin Ma, Xinchao Wang

21 upvotesApril 15, 2025arXiv 预印本
AI 摘要

The survey discusses methods to accelerate reasoning models by compressing Chain-of-Thoughts, developing compact models, and designing efficient decoding strategies.

reasoning modelsChain-of-Thoughts (CoTs)knowledge distillationmodel compression techniquesreinforcement learningefficient decoding strategies

Abstract

Reasoning models have demonstrated remarkable progress in solving complex and logic-intensive tasks by generating extended Chain-of-Thoughts (CoTs) prior to arriving at a final answer. Yet, the emergence of this "slow-thinking" paradigm, with numerous tokens generated in sequence, inevitably introduces substantial computational overhead. To this end, it highlights an urgent need for effective acceleration. This survey aims to provide a comprehensive overview of recent advances in efficient reasoning. It categorizes existing works into three key directions: (1) shorter - compressing lengthy CoTs into concise yet effective reasoning chains; (2) smaller - developing compact language models with strong reasoning capabilities through techniques such as knowledge distillation, other model compression techniques, and reinforcement learning; and (3) faster - designing efficient decoding strategies to accelerate inference. A curated collection of papers discussed in this survey is available in our GitHub repository.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Efficient Reasoning Models: A Survey | TensorX