TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

reinforcement learning 相关论文

201 篇论文 · 按点赞排序

03

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +399 authors

Kimi K3 is a large-scale mixture-of-experts model with native vision and long-context capabilities that improves scaling efficiency and achieves strong performance across coding, reasoning, and agentic tasks.

509Mixture-of-ExpertsKimi Delta AttentionHF ↗arXiv ↗
04

StudentSim: Training LLM-based Student Simulators

Ke Yang, Chenglong Wang, Michel Galley +4 authors

StudentSim trains personalized student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming existing models across chess, writing, and math.

487student simulatorspooled trainingHF ↗arXiv ↗
06

Scaling Automatic Research Agents via World Models

Xiyuan Yang, Sheikh Sarwar, Jingru Cheng +8 authors

World Model RL replaces costly environment execution with a learned world model and applies debiasing and denoising to accelerate post-training of autonomous research agents.

437World Model RLAutoResearchHF ↗arXiv ↗
07

Qwen2.5 Technical Report

Qwen, An Yang, Baosong Yang +39 authors

Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.

380large language modelspre-trainingHF ↗arXiv ↗
10

Group Sequence Policy Optimization

Chujie Zheng, Shixuan Liu, Mingze Li +9 authors

Group Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.

320Group Sequence Policy OptimizationGSPOHF ↗arXiv ↗
13

Agent Learning via Early Experience

Kai Zhang, Xiangchao Chen, Bo Liu +27 authors

Early experience, using agent-generated interaction data without reward signals, improves policy effectiveness and generalization, serving as a bridge between imitation learning and reinforcement learning.

276reinforcement learningearly experienceHF ↗arXiv ↗
15

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +435 authors

ERNIE 5.0 is a production-scale trillion-parameter autoregressive model that unifies multimodal understanding and generation through sparse MoE architecture and elastic training.

269autoregressive foundation modelunified multimodal understandingHF ↗arXiv ↗
16

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu +22 authors

Agentic reinforcement learning transforms large language models into autonomous decision-making agents by leveraging temporally extended POMDPs, enhancing capabilities like planning and reasoning through reinforcement learning.

239agentic reinforcement learningLLM RLHF ↗arXiv ↗
17

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko +22 authors

Kandinsky 5.0 is a family of state-of-the-art generative models for high-resolution images and short videos, featuring model lineups with varying parameters and enhanced training techniques to achieve superior quality and performance.

236foundation modelshigh-resolution image synthesisHF ↗arXiv ↗
1 / 11

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号