TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

supervised fine-tuning 相关论文

86 篇论文 · 按点赞排序

62

Demystifying Long Chain-of-Thought Reasoning in LLMs

Edward Yeo, Yuxuan Tong, Morry Niu +2 authors

Investigation into long chains-of-thought reasoning in large language models reveals the critical role of training compute, reward shaping, and verifiable reward signals in enabling and measuring this capability.

57large language modelslong chains-of-thoughtHF ↗arXiv ↗
64

Kosmos-2.5: A Multimodal Literate Model

Tengchao Lv, Yupan Huang, Jingye Chen +11 authors

Kosmos-2.5, a unified multimodal model, generates spatially-aware and structured text from text-intensive images using a Transformer architecture and task-specific prompts.

56multimodal literate modelmachine readingHF ↗arXiv ↗
67

How to Synthesize Text Data without Model Collapse?

Xuekai Zhu, Daixuan Cheng, Hengli Li +7 authors

The use of synthetic data in language model training leads to model collapse, which is mitigated by token-level editing of human-produced data to create semi-synthetic data.

52synthetic datamodel collapseHF ↗arXiv ↗
68

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

NVIDIA, Alisson Azzolini, Hannah Brandon +42 authors

Cosmos-Reason1 models, using hierarchical and two-dimensional ontologies for physical common sense and embodied reasoning, generate embodied decisions through multimodal large language models trained in vision and Physical AI stages.

52Physical AIreasoningHF ↗arXiv ↗
69

Miles v0.1: Production-Level Post-Training

RadixArk, Tom Chen, Mao Cheng +10 authors

Miles is an open-source, production-ready system for large-scale reinforcement learning and post-training that supports diverse backends, weight synchronization, LoRA, distillation, and diffusion models.

52reinforcement-learningrollout enginesHF ↗arXiv ↗
70

Baichuan Alignment Technical Report

Mingan Lin, Fan Yang, Yanjun Shen +22 authors

Baichuan Alignment provides comprehensive insights into alignment methodologies used in Baichuan models, detailing improvements through Prompt Augmentation System, Supervised Fine-Tuning, and Preference Alignment across various benchmarks.

51Prompt Augmentation SystemSupervised Fine-TuningHF ↗arXiv ↗
73

SemiEvol: Semi-supervised Fine-tuning for LLM Adaptation

Junyu Luo, Xiao Luo, Xiusi Chen +3 authors

A semi-supervised fine-tuning framework named SemiEvol enhances LLM adaptation using both labeled and unlabeled data, showing improved performance through bi-level knowledge propagation and collaborative learning.

47supervised fine-tuninglarge language modelsHF ↗arXiv ↗
77

DiffusionGemma Technical Report

DiffusionGemma Team, Adrien Ali Taïga, James Assiene +41 authors

DiffusionGemma is a fine-tuned mixture-of-experts language model that uses discrete diffusion to generate text blocks in parallel, achieving high speed while preserving capabilities like multimodal inputs and reasoning.

36discrete diffusionmixture-of-expertsHF ↗arXiv ↗
79

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97 authors

InternLM2 is an open-source LLM that outperforms predecessors through innovative pre-training and optimization techniques, including Supervised Fine-Tuning and Conditional Online Reinforcement Learning from Human Feedback.

35Large Language ModelsLLMsHF ↗arXiv ↗
4 / 5

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号