TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 5 – May 11, 2025
本周最热191

Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Andrew Zhao, Yiran Wu, Yang Yue +8 authors

Absolute Zero Reasoner (AZR) achieves state-of-the-art performance on coding and mathematical reasoning tasks through self-generated, verified learning without external data.

reinforcement learning with verifiable rewardszero settingself-evolving training curriculumcode executorHF ↗arXiv ↗

50 篇论文 · 按点赞排序

08

RM-R1: Reward Modeling as Reasoning

Xiusi Chen, Gaotang Li, Ziqi Wang +9 authors

Reasoning Reward Models (ReasRMs) enhance reward modeling for large language models by integrating reasoning tasks, improving interpretability and performance.

81reward modelingreinforcement learning from human feedback (RLHF)HF ↗arXiv ↗
09

Flow-GRPO: Training Flow Matching Models via Online RL

Jie Liu, Gongye Liu, Jiajun Liang +6 authors

Flow-GRPO combines online reinforcement learning with flow matching models through an ODE-to-SDE conversion and denoising reduction, improving sampling efficiency and performance across text-to-image tasks.

77Flow-GRPOonline reinforcement learningHF ↗arXiv ↗
12

Llama-Nemotron: Efficient Reasoning Models

Akhiad Bercovich, Itay Levy, Izik Golan +129 authors

Llama-Nemotron models offer exceptional reasoning capabilities, inference efficiency, and open licensing through neural architecture search, knowledge distillation, and reasoning-focused post-training, including supervised fine-tuning and reinforcement learning.

45neural architecture searchknowledge distillationHF ↗arXiv ↗
18

Practical Efficiency of Muon for Pretraining

Essential AI, Ishaan Shah, Anthony M. Polloreno +21 authors

Muon, a second-order optimizer, improves data efficiency and computational savings over AdamW, especially at large batch sizes, and combined with muP, it provides efficient hyperparameter transfer and minimal resource overhead.

31Muonsecond-order optimizerHF ↗arXiv ↗
27

Scalable Chain of Thoughts via Elastic Reasoning

Yuhui Xu, Hanze Dong, Lei Wang +3 authors

Elastic Reasoning is a framework that divides reasoning into thinking and solution phases with separate budgets, improving model reliability and efficiency under resource constraints.

26Elastic Reasoningchain of thoughtsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号