TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

fine-tuning 相关论文

75 篇论文 · 按点赞排序

45

Style-Friendly SNR Sampler for Style-Driven Generation

Jooyoung Choi, Chaehun Shin, Yeongtak Oh +2 authors

The Style-friendly SNR sampler modifies the noise level distribution during fine-tuning to improve style alignment in diffusion models, enabling better capture of unique artistic styles.

40diffusion modelssignal-to-noise ratio (SNR)HF ↗arXiv ↗
47

FuseChat: Knowledge Fusion of Chat Models

Fanqi Wan, Ziyi Yang, Longguang Zhong +3 authors

FuseChat extends the FuseLLM framework for knowledge fusion of chat LLMs through lightweight fine-tuning and parameter merging, achieving superior performance across various domains.

39knowledge fusionfine-tuningHF ↗arXiv ↗
49

JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Lianghui Zhu, Xinggang Wang, Xinlong Wang

Large Language Models fine-tuned as scalable judges (JudgeLM) achieve state-of-the-art performance in evaluating open-ended benchmarks through a comprehensive dataset and benchmark, enhancing judgment efficiency and accuracy.

35Large Language Modelsfine-tuningHF ↗arXiv ↗
53

Can large language models explore in-context?

Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster +2 authors

Experimentation with LLMs in multi-armed bandit environments reveals that robust exploration requires specific prompts, external summarization of history, or algorithmic interventions.

33Large Language ModelsLLMsHF ↗arXiv ↗
54

FP8-LM: Training FP8 Large Language Models

Houwen Peng, Kan Wu, Yixuan Wei +17 authors

A new FP8 automatic mixed-precision framework for training large language models reduces memory usage and increases speed compared to BF16 and Nvidia Transformer Engine.

33FP8low-bit data formatsHF ↗arXiv ↗
57

Octo: An Open-Source Generalist Robot Policy

Octo Model Team, Dibya Ghosh, Homer Walke +15 authors

Octo, a large transformer-based policy trained on extensive robotic datasets, demonstrates versatility and efficient fine-tuning for diverse robotic platforms and tasks.

29transformer-based policyOpen X-Embodiment datasetHF ↗arXiv ↗
60

Self-Play Preference Optimization for Language Model Alignment

Yue Wu, Zhiqing Sun, Huizhuo Yuan +3 authors

A self-play method called SPPO for language model alignment achieves state-of-the-art performance by approximating Nash equilibrium policy in a constant-sum game setting, outperforming other approaches with limited data.

29reinforcement learning from human feedbackparametric modelsHF ↗arXiv ↗
3 / 4

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号