TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

reinforcement learning 相关论文

201 篇论文 · 按点赞排序

141

LLaDA2.1: Speeding Up Text Diffusion via Token Editing

Tiwei Bie, Maosong Cao, Xiang Cao +47 authors

LLaDA2.1 introduces a novel token-to-token editing approach with speed and quality modes, enhanced through reinforcement learning for improved reasoning and instruction following in large language diffusion models.

73block-diffusion modelsdecoding speedHF ↗arXiv ↗
143

Magistral

Mistral-AI, Abhinav Rastogi, Albert Q. Jiang +97 authors

Magistral, a scalable reinforcement learning pipeline, demonstrates that RL can enhance multimodal understanding and instruction following in large language models without requiring existing RL traces.

69reinforcement learningRLHF ↗arXiv ↗
144

Competitive Programming with Large Reasoning Models

OpenAI, Ahmed El-Kishky, Alexander Wei +22 authors

General-purpose reinforcement learning applied to large language models outperforms domain-specific systems in complex coding and reasoning tasks, achieving top results in competitions without hand-crafted strategies.

69reinforcement learninglarge language modelsHF ↗arXiv ↗
145

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs

Baorong Shi, Bo Cui, Boyuan Jiang +17 authors

MedXIAOHE is a medical vision-language foundation model that enhances clinical understanding through entity-aware continual pretraining, reinforcement learning, and tool-augmented agentic training for reliable diagnostic reasoning.

68vision-language foundation modelentity-aware continual pretrainingHF ↗arXiv ↗
146

π_RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

Kang Chen, Zhihao Liu, Tonghe Zhang +10 authors

The framework π<sub>RL</sub> uses reinforcement learning to train flow-based Vision-Language-Action models, addressing challenges with intractable action log-likelihoods and achieving significant performance improvements over supervised fine-tuning.

66reinforcement learningsupervised fine-tuningHF ↗arXiv ↗
149

Controllable Text Generation for Large Language Models: A Survey

Xun Liang, Hanyu Wang, Yezhaohui Wang +8 authors

Controllable Text Generation techniques for Large Language Models ensure predefined control conditions and high-quality text output, covering content and attribute control through various methods like retraining, fine-tuning, and latent manipulation.

65Large Language ModelsControllable Text GenerationHF ↗arXiv ↗
151

One RL to See Them All: Visual Triple Unified Reinforcement Learning

Yan Ma, Linge Du, Xuyang Shen +7 authors

A unified reinforcement learning system, V-Triune, combines visual reasoning and perception tasks in vision-language models through a single training pipeline, achieving significant improvements across various tasks.

63visual triple unified reinforcement learningsample-level data formattingHF ↗arXiv ↗
153

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Jiazhan Feng, Shijue Huang, Xingwei Qu +6 authors

ReTool, a tool-integrated learning framework, enhances reasoning models with real-time code execution and reinforcement learning, significantly improving performance in structured problem-solving tasks like mathematical reasoning.

63reasoning modelsreinforcement learningHF ↗arXiv ↗
154

Iris: Climbing to the Search Frontier

Ziyuan Liu, Hengqi Liu, Zichuan Wang +6 authors

Two large-scale search agents are trained via a multi-stage pipeline combining supervised fine-tuning and reinforcement learning against live search, achieving state-of-the-art open-source results on complex web benchmarks through rigorous trajectory filtering and inference-time context management.

62search agentsmulti-hop chainsHF ↗arXiv ↗
156

Process Reinforcement through Implicit Rewards

Ganqu Cui, Lifan Yuan, Zefan Wang +20 authors

PRIME leverages implicit process rewards to improve the reinforcement learning of large language models, achieving better performance with less data compared to traditional methods.

62dense process rewardssparse outcome-level rewardsHF ↗arXiv ↗
8 / 11

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号