Token-Budget-Aware LLM Reasoning
Tingxu Han, Chunrong Fang, Shiyu Zhao +3 authors
A token-budget-aware framework dynamically estimates and allocates token budgets for LLM reasoning, reducing costs with minimal performance loss.
Trends · 研究趋势
数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。
Tingxu Han, Chunrong Fang, Shiyu Zhao +3 authors
A token-budget-aware framework dynamically estimates and allocates token budgets for LLM reasoning, reducing costs with minimal performance loss.
Jian Hu, Xibin Wu, Weixun Wang +3 authors
OpenRLHF is an open-source RLHF framework that enhances training efficiency and accessibility for researchers and practitioners.
Zelong Sun, Jun Wang, Kaicheng Yang +3 authors
UniME-R1 improves multimodal retrieval by generating retrieval-centric reasoning guided by initial candidate feedback rather than query-only explanations.
Sean O'Brien, Mike Lewis
Contrastive Decoding, a simple and training-free text generation method, outperforms other decoding techniques on various reasoning tasks and logical benchmarks by improving perceived quality and reducing reasoning errors.
Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez +7 authors
Chain-of-thought prompting enhances large language model performance primarily in math or logic tasks by improving symbolic reasoning, but it is less beneficial in other domains and can be selectively applied to save inference costs.
Pei Zhou, Aman Madaan, Srividya Pranavi Potharaju +9 authors
LLMs perform well in inferring beliefs but struggle translating these into actions; Foresee and Reflect (FaR) improves their action guidance in social scenarios.
Jonghyun Song, Sangjun Song, Minjae Oh +3 authors
SHAPE analyzes chain-of-thought reasoning via semantic spaces and heuristics to diagnose LLM mathematical reasoning and improve post-training.
Guillaume Sanchez, Honglu Fan, Alexander Spangher +3 authors
Classifier-Free Guidance enhances performance across various language modeling tasks and improves the faithfulness and coherence of AI assistants, outperforming models with higher parameter counts.
Tamera Lanham, Anna Chen, Ansh Radhakrishnan +27 authors
Larger language models may not produce faithful reasoning when using chain-of-thought despite performance improvements, depending on model size and task.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号