TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

transformer 相关论文

22 篇论文 · 按点赞排序

02

Differential Transformer

Tianzhu Ye, Li Dong, Yuqing Xia +4 authors

Diff Transformer improves large language models by selectively focusing attention on relevant context and reducing noise, leading to better performance in scaling, long-context modeling, key information retrieval, and in-context learning.

183TransformerDiff TransformerHF ↗arXiv ↗
06

The Llama 3 Herd of Models

Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey +530 authors

Llama 3, a multilingual and multi-modal language model with 405B parameters, achieves competitive performance across tasks including image, video, and speech recognition when integrated through a compositional approach.

119TransformermultilingualityHF ↗arXiv ↗
08

Movie Gen: A Cast of Media Foundation Models

Adam Polyak, Amit Zohar, Andrew Brown +85 authors

Movie Gen, a suite of foundation models, generates high-quality videos with synchronized audio, excelling in various tasks through architectural, training, and technical innovations.

100transformervideo tokensHF ↗arXiv ↗
09

Emu3: Next-Token Prediction is All You Need

Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22 authors

Emu3, a transformer-based multimodal model trained exclusively with next-token prediction, outperforms existing diffusion and compositional models in generation and perception tasks.

99next-token predictionmultimodal modelsHF ↗arXiv ↗
10

Next-Embedding Prediction Makes Strong Vision Learners

Sihan Xu, Ziqiao Ma, Wenhao Chai +5 authors

Generative pretraining using next embedding prediction outperforms traditional self-supervised methods in visual learning tasks, achieving high accuracy on ImageNet and effective transfer to semantic segmentation.

91generative pretrainingpredictive tasksHF ↗arXiv ↗
17

In-Context Learning Creates Task Vectors

Roee Hendel, Mor Geva, Amir Globerson

In-Context Learning in Large Language Models can be understood as compressing a training set into a task vector that modulates a transformer for output generation.

43in-context learninglarge language modelsHF ↗arXiv ↗
19

Wan-Streamer v0.2: Higher Resolution, Same Latency

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23 authors

Wan-Streamer v0.2 enhances audio-visual interaction by increasing visual resolution while maintaining low latency through optimized thinker-performer architecture with multi-GPU parallel processing.

37streaming perceptionTransformerHF ↗arXiv ↗
20

RMT: Retentive Networks Meet Vision Transformers

Qihang Fan, Huaibo Huang, Mingrui Chen +2 authors

The proposed RMT model, combining RetNet and Transformer architectures, introduces explicit spatial distance priors and coordinate-wise decomposition to achieve exceptional performance in computer vision tasks.

34TransformerRetentive Network (RetNet)HF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号