TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

reinforcement learning 相关论文

201 篇论文 · 按点赞排序

181

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Yifan Shen, Jian Xu, Boyi Li +6 authors

ChronoVision improves visual temporal reasoning by aligning latent imagery with logic through reconstructive prediction, ROI attention, and reinforcement learning, achieving strong results on video reasoning benchmarks.

39multimodal large language modelsReconstructive Visual HeadHF ↗arXiv ↗
185

DiffusionGemma Technical Report

DiffusionGemma Team, Adrien Ali Taïga, James Assiene +41 authors

DiffusionGemma is a fine-tuned mixture-of-experts language model that uses discrete diffusion to generate text blocks in parallel, achieving high speed while preserving capabilities like multimodal inputs and reasoning.

36discrete diffusionmixture-of-expertsHF ↗arXiv ↗
188

Can large language models explore in-context?

Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster +2 authors

Experimentation with LLMs in multi-armed bandit environments reveals that robust exploration requires specific prompts, external summarization of history, or algorithmic interventions.

33Large Language ModelsLLMsHF ↗arXiv ↗
189

SHAPE of Chain-of-Thought in Math Reasoning

Jonghyun Song, Sangjun Song, Minjae Oh +3 authors

SHAPE analyzes chain-of-thought reasoning via semantic spaces and heuristics to diagnose LLM mathematical reasoning and improve post-training.

31Chain-of-Thoughtsemantic spacesHF ↗arXiv ↗
194

LIMA: Less Is More for Alignment

Chunting Zhou, Pengfei Liu, Puxin Xu +12 authors

A 65B parameter LLaMa language model trained with minimal instructional data matches or outperforms models with extensive human preference modeling in most cases, indicating that pretraining is predominantly responsible for knowledge acquisition.

27large language modelsunsupervised pretrainingHF ↗arXiv ↗
198

Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Mihir Prabhudesai, Anirudh Goyal, Deepak Pathak +1 authors

AlignProp refines text-to-image diffusion models using backpropagation through the denoising process, leveraging low-rank adapters and gradient checkpointing to optimize for various objectives with higher efficiency.

22text-to-image diffusion modelsreinforcement learningHF ↗arXiv ↗
200

Xiaomi-GUI-0 Technical Report

Wanxia Cao, Chengzhen Duan, Pei Fu +27 authors

A native multimodal GUI agent trained in real-device environments demonstrates superior performance and stability compared to traditional benchmark-based approaches.

19vision-language modelsinterface actionsHF ↗arXiv ↗
10 / 11

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号