TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

supervised fine-tuning 相关论文

86 篇论文 · 按点赞排序

05

Recursive Synthesis for Long-Horizon Terminal Tasks

Zhongzhi Li, Yucheng Shi, Zongxia Li +8 authors

Recursive verified synthesis generates scalable long-horizon terminal-agent training data, substantially improving model performance on terminal benchmarks through supervised fine-tuning and reinforcement learning.

252recursive verified synthesisterminal-agent tasksHF ↗arXiv ↗
06

Agentic Reasoning for Large Language Models

Tianxin Wei, Ting-Wei Li, Zhining Liu +26 authors

Agentic reasoning redefines large language models as autonomous agents capable of planning, acting, and learning through continuous interaction in dynamic environments across single-agent and multi-agent frameworks.

207large language modelsagentic reasoningHF ↗arXiv ↗
09

Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling

Yafu Li, Runzhe Zhan, Haoran Zhang +25 authors

A systematic approach transforms post-trained reasoning models into rigorous olympiad-level solvers through reverse-perplexity curriculum, two-stage reinforcement learning, and test-time scaling, achieving gold-medal performance on mathematical and physics competitions.

165reasoning modelsmathematical problem solvingHF ↗arXiv ↗
13

Large Language Diffusion Models

Shen Nie, Fengqi Zhu, Zebin You +7 authors

LLaDA, a diffusion model trained from scratch, outperforms autoregressive models in benchmarks and demonstrates strong instruction-following capabilities, challenging the dominance of ARMs in LLMs.

128autoregressive modelsLLaDAHF ↗arXiv ↗
17

The Differences Between Direct Alignment Algorithms are a Blur

Alexey Gorbatovski, Boris Shaposhnikov, Viacheslav Sinii +2 authors

Direct Alignment Algorithms improve language model alignment by introducing a supervised fine-tuning phase and adjusting preference optimization strength, showing that ranking objectives are crucial for performance.

113Direct Alignment AlgorithmsRLHFHF ↗arXiv ↗
18

VIDEOP2R: Video Understanding from Perception to Reasoning

Yifan Jiang, Yueying Wang, Rui Zhao +4 authors

VideoP2R, a process-aware reinforcement fine-tuning framework, improves video reasoning and understanding by modeling perception and reasoning separately, achieving state-of-the-art results on multiple benchmarks.

113Reinforcement fine-tuningsupervised fine-tuningHF ↗arXiv ↗
19

Grounding Computer Use Agents on Human Demonstrations

Aarash Feizi, Shravan Nayak, Xiangru Jian +14 authors

GroundCUA, a large-scale desktop grounding dataset, enables the development of GroundNext models that achieve state-of-the-art performance in mapping instructions to UI elements with less training data.

107groundingnatural language instructionsHF ↗arXiv ↗
20

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Tong Zheng, Hongming Zhang, Wenhao Yu +8 authors

Parallel-R1, a reinforcement learning framework, enhances large language models' reasoning capabilities by enabling parallel thinking through a progressive curriculum, leading to significant performance improvements on math benchmarks.

105parallel thinkingreinforcement learningHF ↗arXiv ↗
1 / 5

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
11 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
13
37 篇论文
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号