TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

reinforcement learning 相关论文

201 篇论文 · 按点赞排序

162

Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance

Lingfei Qian, Weipeng Zhou, Yan Wang +3 authors

A study evaluates 16 large language models on complex financial tasks, finding that domain-specific CoT fine-tuning and reinforcement learning improve performance and highlight the need for further research on long-context and multi-table reasoning.

59large language modelsfinancial reasoningHF ↗arXiv ↗
166

Demystifying Long Chain-of-Thought Reasoning in LLMs

Edward Yeo, Yuxuan Tong, Morry Niu +2 authors

Investigation into long chains-of-thought reasoning in large language models reveals the critical role of training compute, reward shaping, and verifiable reward signals in enabling and measuring this capability.

57large language modelslong chains-of-thoughtHF ↗arXiv ↗
170

Transformer^2: Self-adaptive LLMs

Qi Sun, Edoardo Cetin, Yujin Tang

A self-adaptive framework for large language models uses reinforcement learning to dynamically adjust task-specific components during inference, enhancing adaptability and performance with efficiency.

55self-adaptive large language models (LLMs)fine-tuningHF ↗arXiv ↗
174

Normalized Low-Rank Adaptation

Jiale Kang, Ziyin Yue, Zheng Zhan +2 authors

Normalized Low-Rank Adaptation stabilizes LoRA training by normalizing down-projection matrices, accelerating convergence and improving performance without extra parameters or inference cost.

52low-rank adaptationLoRAHF ↗arXiv ↗
176

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Junyao Yang, Yucheng Shi, Zhongzhi Li +4 authors

T1 is a 122B Mixture-of-Experts model trained with reinforcement learning to execute long-horizon terminal tasks in a cloud sandbox, achieving state-of-the-art results through stable actor-critic optimization and out-of-distribution training.

46Mixture-of-Expertsreinforcement learningHF ↗arXiv ↗
180

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Philip Anastassiou, Jiawei Chen, Jitong Chen +43 authors

Seed-TTS is a family of large-scale TTS models that generate high-quality speech with in-context learning, superior controllability, and a non-autoregressive variant using diffusion-based architecture that does not rely on pre-estimated phoneme durations.

39autoregressive text-to-speechspeech generationHF ↗arXiv ↗
9 / 11

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号