TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

transformers 相关论文

23 篇论文 · 按点赞排序

01

HRM-Text: Efficient Pretraining Beyond Scaling

Guan Wang, Changling Liu, Chenyu Wang +6 authors

A Hierarchical Recurrent Model architecture with specialized training on instruction-response pairs achieves competitive language modeling performance with significantly reduced computational requirements compared to traditional Transformer-based approaches.

321Hierarchical Recurrent ModelTransformersHF ↗arXiv ↗
03

Transformers without Normalization

Jiachen Zhu, Xinlei Chen, Kaiming He +2 authors

Dynamic Tanh (DyT) replaces normalization layers in Transformers, achieving equivalent or superior performance without hyperparameter tuning across various tasks.

172Normalization layersDynamic TanhHF ↗arXiv ↗
04

RWKV-7 "Goose" with Expressive Dynamic State Evolution

Bo Peng, Ruichong Zhang, Daniel Goldstein +12 authors

RWKV-7 "Goose" achieves state-of-the-art performance in multilingual tasks with optimal memory and inference efficiency, exceeding Transformer capabilities in complexity.

153sequence modeling architecturedelta ruleHF ↗arXiv ↗
06

Vision Transformers Need Registers

Timothée Darcet, Maxime Oquab, Julien Mairal +1 authors

Additional input tokens in Vision Transformers mitigate artifacts in feature maps, enhancing performance and enabling smoother visual processing.

86Transformersfeature mapsHF ↗arXiv ↗
11

Transformers Can Do Arithmetic with the Right Embeddings

Sean McLeish, Arpit Bansal, Alex Stein +8 authors

Transformers achieve state-of-the-art performance on large arithmetic tasks and other reasoning tasks by addressing positional tracking with embeddings and integrating architectural modifications.

55transformerspositional trackingHF ↗arXiv ↗
14

Transformers meet Neural Algorithmic Reasoners

Wilfried Bounsi, Borja Ibarz, Andrew Dudzik +5 authors

A novel TransNAR model combines Transformer-based language understanding with graph neural network solvers to enhance algorithmic reasoning.

44Transformersnatural language understandingHF ↗arXiv ↗
15

TransformerFAM: Feedback attention is working memory

Dongseong Hwang, Weiran Wang, Zhuoyuan Huo +2 authors

Feedback Attention Memory (FAM) enhances Transformer architecture by enabling long-context processing without additional weights, significantly improving performance on large sequences across various model sizes.

43TransformersFeedback Attention MemoryHF ↗arXiv ↗
19

Transformers are Multi-State RNNs

Matanel Oren, Michael Hassid, Yossi Adi +1 authors

Decoder-only transformers can be conceptualized as finite multi-state RNNs, and a new cache compression technique, TOVA, significantly reduces their computational cost while maintaining high performance.

39transformersrecurrent neural networksHF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号