JudgeLRM: Large Reasoning Models as a Judge
Nuo Chen, Zhiyuan Hu, Qingyun Zou +4 authors
JudgeLRM models trained with reinforcement learning outperform existing models, especially in tasks requiring deep reasoning.
Trends · 研究趋势
数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。
Nuo Chen, Zhiyuan Hu, Qingyun Zou +4 authors
JudgeLRM models trained with reinforcement learning outperform existing models, especially in tasks requiring deep reasoning.
Yafu Li, Xuyang Hu, Xiaoye Qu +2 authors
TPO refines LLM outputs in real-time using textual critiques, achieving preference alignment without retraining.
Tianle Cai, Yuhong Li, Zhengyang Geng +4 authors
Medusa enhances Large Language Model inference by adding parallel decoding heads and using tree-based attention to predict multiple tokens simultaneously, achieving significant speedup with minimal latency.
Liang Chen, Zekun Wang, Shuhuai Ren +24 authors
A comprehensive taxonomy unifies multimodal understanding and generation through Next Token Prediction, covering aspects like tokenization, model architectures, task representation, datasets, and open challenges.
Christopher E. Mower, Yuhui Wan, Hongzhan Yu +19 authors
A framework enables non-experts to program robots using natural language prompts and contextual information from the Robot Operating System (ROS), supported by large language models and offering various behavior modes, imitation learning, and feedback mechanisms.
Ryan Teknium, Jeffrey Quesnelle, Chen Guang
Hermes 3, a neutrally-aligned instruct and tool use model with strong reasoning and creative capabilities, achieves top performance on public benchmarks.
Daniil Laptev, Nikita Balagansky, Yaroslav Aksenov +1 authors
A new data-free cosine similarity method maps and steers features across layers of large language models, enhancing interpretability and targeted control in text generation.
Lingfei Qian, Weipeng Zhou, Yan Wang +3 authors
A study evaluates 16 large language models on complex financial tasks, finding that domain-specific CoT fine-tuning and reinforcement learning improve performance and highlight the need for further research on long-context and multi-table reasoning.
Junlin Wang, Jue Wang, Ben Athiwaratkun +2 authors
A Mixture-of-Agents (MoA) methodology combines multiple large language models (LLMs) to achieve state-of-the-art performance on various evaluations.
Hengguang Zhou, Xirui Li, Ruochen Wang +3 authors
A non-SFT model replicated emergent reasoning characteristics for multimodal tasks using reinforcement learning, achieving higher accuracy than base and SFT models on CVBench.
Edward Yeo, Yuxuan Tong, Morry Niu +2 authors
Investigation into long chains-of-thought reasoning in large language models reveals the critical role of training compute, reward shaping, and verifiable reward signals in enabling and measuring this capability.
Junyou Li, Qin Zhang, Yangbin Yu +2 authors
A sampling-and-voting method enhances large language models' performance by increasing the number of agents, with effectiveness tied to task difficulty.
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger +25 authors
LLM360 initiative promotes full transparency and reproducibility in LLM training by open-sourcing training code, data, model checkpoints, and intermediate results.
Yuhao Wu, Yushi Bai, Zhiqiang Hu +2 authors
An incentivization-based reinforcement learning approach is used to develop a large language model capable of generating ultra-long, high-quality text without the need for synthetic data or supervised fine-tuning.
Wenxuan Zhang, Hou Pong Chan, Yiran Zhao +9 authors
SeaLLMs 3 is a large language model optimized for Southeast Asian languages, offering state-of-the-art performance in various tasks while addressing safety and cultural considerations.
Chaofan Tao, Qian Liu, Longxu Dou +5 authors
Investigating the impact of vocabulary size on the scaling of large language models reveals that larger vocabularies improve performance when considering compute budgets, and current models often use suboptimal vocabulary sizes.
Guosheng Dong, Da Pan, Yiding Sun +17 authors
A data processing pipeline and open-sourced details for training a large language model achieve competitive performance with commercial models.
Praveen K Kanithi, Clément Christophe, Marco AF Pimentel +7 authors
MEDIC framework evaluates Large Language Models across five clinical dimensions to guide model selection in healthcare applications, identifying performance trade-offs and ensuring practical implementation.
Zorik Gekhman, Eyal Ben David, Hadas Orgad +5 authors
LLMs encode more internal factual knowledge than they express externally, with some knowledge so deeply hidden that it is never generated, despite repeated sampling.
Ye Tian, Baolin Peng, Linfeng Song +4 authors
AlphaLLM enhances Large Language Models through integration with Monte Carlo Tree Search and critic models, improving performance in complex reasoning tasks without additional annotations.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号