TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

pre-training 相关论文

22 篇论文 · 按点赞排序

01

Qwen2.5 Technical Report

Qwen, An Yang, Baosong Yang +39 authors

Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.

380large language modelspre-trainingHF ↗arXiv ↗
03

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko +22 authors

Kandinsky 5.0 is a family of state-of-the-art generative models for high-resolution images and short videos, featuring model lineups with varying parameters and enhanced training techniques to achieve superior quality and performance.

236foundation modelshigh-resolution image synthesisHF ↗arXiv ↗
06

Large Language Diffusion Models

Shen Nie, Fengqi Zhu, Zebin You +7 authors

LLaDA, a diffusion model trained from scratch, outperforms autoregressive models in benchmarks and demonstrates strong instruction-following capabilities, challenging the dominance of ARMs in LLMs.

128autoregressive modelsLLaDAHF ↗arXiv ↗
11

LIMO: Less is More for Reasoning

Yixin Ye, Zhen Huang, Yang Xiao +3 authors

LIMO, a new model, achieves high mathematical reasoning performance using minimal training data, challenging the notion that extensive datasets are necessary for complex reasoning.

62LIMOLIMO HypothesisHF ↗arXiv ↗
13

LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Fuzhao Xue, Yukang Chen, Dacheng Li +15 authors

LongVILA, a full-stack solution for long-context vision-language models, introduces Multi-Modal Sequence Parallelism for efficient training and inference, and a five-stage training pipeline that enhances long video processing and captioning performance.

52Multi-Modal Sequence ParallelismMM-SPHF ↗arXiv ↗
14

How to Synthesize Text Data without Model Collapse?

Xuekai Zhu, Daixuan Cheng, Hengli Li +7 authors

The use of synthetic data in language model training leads to model collapse, which is mitigated by token-level editing of human-produced data to create semi-synthetic data.

52synthetic datamodel collapseHF ↗arXiv ↗
15

Weaver: Foundation Models for Creative Writing

Tiannan Wang, Jiamin Chen, Qingrui Jia +43 authors

Weaver, a family of specialized large language models, achieves superior writing capabilities through pre-training and fine-tuning methods, surpasses GPT-4 in various writing tasks, and supports retrieval-augmented generation and tool usage.

46large language modelspre-trainingHF ↗arXiv ↗
17

How to Train Data-Efficient LLMs

Noveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang +6 authors

Data-efficient methods like Ask-LLM and Density sampling improve model quality and training efficiency in large language models by optimizing data selection and coverage.

43large language modelsdata-efficientHF ↗arXiv ↗
18

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97 authors

InternLM2 is an open-source LLM that outperforms predecessors through innovative pre-training and optimization techniques, including Supervised Fine-Tuning and Conditional Online Reinforcement Learning from Human Feedback.

35Large Language ModelsLLMsHF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
11 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
13
37 篇论文
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号