TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

instruction-following 相关论文

10 篇论文 · 按点赞排序

01

Large Language Diffusion Models

Shen Nie, Fengqi Zhu, Zebin You +7 authors

LLaDA, a diffusion model trained from scratch, outperforms autoregressive models in benchmarks and demonstrates strong instruction-following capabilities, challenging the dominance of ARMs in LLMs.

128autoregressive modelsLLaDAHF ↗arXiv ↗
02

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Ruiyan Gong, Yingnan Guo, Junjun Hu +37 authors

ABot-N1 improves visual language navigation by separating reasoning from control through a slow-fast architecture that uses explicit chain-of-thought reasoning and pixel goals to guide continuous waypoint generation, achieving strong results across diverse embodied tasks.

102Visual Language NavigationChain-of-Thought reasoningHF ↗arXiv ↗
03

Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Keen You, Haotian Zhang, Eldon Schoop +5 authors

Ferret-UI, a multimodal large language model tailored for mobile UI screens, enhances understanding and interaction through region annotations and a comprehensive dataset of UI tasks, outperforming existing models including GPT-4V.

83multimodal large language modelsMLLMsHF ↗arXiv ↗
04

LLaDA2.1: Speeding Up Text Diffusion via Token Editing

Tiwei Bie, Maosong Cao, Xiang Cao +47 authors

LLaDA2.1 introduces a novel token-to-token editing approach with speed and quality modes, enhanced through reinforcement learning for improved reasoning and instruction following in large language diffusion models.

73block-diffusion modelsdecoding speedHF ↗arXiv ↗
06

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +213 authors

Gemma 3 introduces vision capabilities, broader language coverage, and extended context length, featuring an optimized architecture and post-training enhancements to outperform previous versions.

58multimodal modelsvision understandingHF ↗arXiv ↗
07

LLaMA Pro: Progressive LLaMA with Block Expansion

Chengyue Wu, Yukang Gan, Yixiao Ge +5 authors

A new post-pretraining method using expanded Transformer blocks for Large Language Models improves knowledge without catastrophic forgetting, yielding LLaMA Pro-8.3B that excels in general tasks, programming, and mathematics.

54Large Language ModelsLLMsHF ↗arXiv ↗
08

VeRA: Vector-based Random Matrix Adaptation

Dawid Jan Kopiczko, Tijmen Blankevoort, Yuki Markus Asano

Vector-based Random Matrix Adaptation (VeRA) reduces the number of trainable parameters by 10x compared to LoRA while maintaining performance, and is demonstrated on benchmarks like GLUE and E2E, showing its utility in instruction-following.

30Low-rank adaptationLoRAHF ↗arXiv ↗
09

Knowledge Distillation of Large Language Models

Yuxian Gu, Li Dong, Furu Wei +1 authors

MiniLLM distills knowledge from large generative language models to smaller models using reverse KLD for better precision, quality, and performance.

23Knowledge Distillation (KD)large language models (LLMs)HF ↗arXiv ↗

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
11 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
13
37 篇论文
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号