TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

reinforcement learning 相关论文

201 篇论文 · 按点赞排序

23

Agentic Reasoning for Large Language Models

Tianxin Wei, Ting-Wei Li, Zhining Liu +26 authors

Agentic reasoning redefines large language models as autonomous agents capable of planning, acting, and learning through continuous interaction in dynamic environments across single-agent and multi-agent frameworks.

207large language modelsagentic reasoningHF ↗arXiv ↗
25

STEP3-VL-10B Technical Report

Ailin Huang, Chengyuan Yao, Chunrui Han +90 authors

STEP3-VL-10B achieves superior multimodal performance through unified pre-training with a language-aligned Perception Encoder and Qwen3-8B decoder, combined with scaled post-training and Parallel Coordinated Reasoning for efficient large-scale visual reasoning.

196multimodal tokensPerception EncoderHF ↗arXiv ↗
32

General Agentic Memory Via Deep Research

B. Y. Yan, Chaofan Li, Hongjin Qian +2 authors

GAM, a novel framework that employs JIT compilation principles, improves memory efficiency and task completion by leveraging a lightweight memorizer and researcher in conjunction with reinforcement learning.

172general agentic memoryGAMHF ↗arXiv ↗
33

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Jing Liang, Hongyao Tang, Yi Ma +9 authors

Training-inference mismatch in reinforcement learning for large language models leads to instability, which is addressed through a new policy optimization objective and framework that ensures consistent policy improvements between training and inference phases.

170reinforcement learninglarge language modelsHF ↗arXiv ↗
34

Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling

Yafu Li, Runzhe Zhan, Haoran Zhang +25 authors

A systematic approach transforms post-trained reasoning models into rigorous olympiad-level solvers through reverse-perplexity curriculum, two-stage reinforcement learning, and test-time scaling, achieving gold-medal performance on mathematical and physics competitions.

165reasoning modelsmathematical problem solvingHF ↗arXiv ↗
35

Agentic Reinforced Policy Optimization

Guanting Dong, Hangyu Mao, Kai Ma +11 authors

Agentic Reinforced Policy Optimization (ARPO) enhances multi-turn reasoning in large language models by balancing long-horizon capabilities and tool interactions, using entropy-based adaptive rollouts and advantage attribution.

161reinforcement learningverifiable rewardsHF ↗arXiv ↗
36

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi +11 authors

A framework scales vision-language models for long video reasoning using reinforcement learning, achieving strong performance on benchmarks and demonstrating consistent gains with increased video frames.

161vision-language modelsreinforcement learningHF ↗arXiv ↗
37

OpenClaw-RL: Train Any Agent Simply by Talking

Yinjie Wang, Xuyang Chen, Xiaolong Jin +2 authors

OpenClaw-RL framework enables policy learning from diverse next-state signals across multiple interaction modalities using asynchronous training with PRM judges and hindsight-guided distillation.

158agentic RLnext-state signalsHF ↗arXiv ↗
2 / 11

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号