TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

reinforcement learning with verifiable rewards 相关论文

17 篇论文 · 按点赞排序

03

Self-Distilled RLVR

Chenxu Yang, Chuanyu Qin, Qingyi Si +7 authors

RLSD combines reinforcement learning with verifiable rewards and self-distillation to achieve stable training with fine-grained updates and reliable policy direction from environmental feedback.

179on-policy distillationon-policy self-distillationHF ↗arXiv ↗
06

Quantile Advantage Estimation for Entropy-Safe Reasoning

Junkang Wu, Kexin Huang, Jiancan Wu +3 authors

Quantile Advantage Estimation stabilizes reinforcement learning with verifiable rewards by addressing entropy issues and improving performance on large language models.

119Reinforcement Learning with Verifiable Rewardsentropy collapseHF ↗arXiv ↗
07

Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation

Yanqi Dai, Yuxiang Ji, Xiao Zhang +3 authors

MathForge enhances mathematical reasoning in large models through a dual framework combining difficulty-aware policy optimization and multi-aspect question reformulation to address limitations in existing reinforcement learning methods.

119Reinforcement Learning with Verifiable RewardsGroup Relative Policy OptimizationHF ↗arXiv ↗
08

Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text

Ximing Lu, David Acuna, Jaehun Jung +12 authors

Golden Goose synthesizes unlimited RLVR tasks from unverifiable internet text by creating multiple-choice question-answering versions of fill-in-the-middle tasks, enabling large-scale training and achieving state-of-the-art results in cybersecurity and other domains.

113Reinforcement Learning with Verifiable RewardsLarge Language ModelsHF ↗arXiv ↗
12

The Invisible Leash: Why RLVR May Not Escape Its Origin

Fang Wu, Weihao Xuan, Ximing Lu +2 authors

Reinforcement Learning with Verifiable Rewards (RLVR) enhances precision but may limit exploration and discovery of new solutions, suggesting potential limits to its effectiveness in expanding reasoning capabilities.

85Reinforcement Learning with Verifiable RewardsRLVRHF ↗arXiv ↗
13

VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use

Dongfu Jiang, Yi Lu, Zhuofeng Li +9 authors

VerlTool is a unified and modular framework for Agentic Reinforcement Learning with Tool use, addressing inefficiencies in existing approaches and providing competitive performance across multiple domains.

81Reinforcement Learning with Verifiable RewardsAgentic Reinforcement Learning with Tool useHF ↗arXiv ↗
14

RM-R1: Reward Modeling as Reasoning

Xiusi Chen, Gaotang Li, Ziqi Wang +9 authors

Reasoning Reward Models (ReasRMs) enhance reward modeling for large language models by integrating reasoning tasks, improving interpretability and performance.

81reward modelingreinforcement learning from human feedback (RLHF)HF ↗arXiv ↗
17

Cliff: Learning Process Rewards from the First Mistake

Peixuan Han, Runhui Wang, Ketan Ramaneti +3 authors

Cliff improves reinforcement learning with verifiable rewards by using an off-the-shelf language model to detect the first reasoning error and shaping token-level advantages accordingly.

15reinforcement learning with verifiable rewardsprocess reward modelingHF ↗arXiv ↗

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号