TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

large language models 相关论文

319 篇论文 · 按点赞排序

221

Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM

Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma +8 authors

Branch-Train-MiX (BTX) method enhances Large Language Models by asynchronously training experts in parallel and integrating them using Mixture-of-Expert layers with token-level routing for improved accuracy and efficiency.

45Large Language ModelsBranch-Train-MiXHF ↗arXiv ↗
222

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Xiwei Hu, Rui Wang, Yixiao Fang +3 authors

ELLA, an Efficient Large Language Model Adapter, enhances text-to-image diffusion models by integrating powerful Large Language Models through a Timestep-Aware Semantic Connector, improving dense prompt comprehension and generation quality.

45diffusion modelstext-to-image generationHF ↗arXiv ↗
225

Foundation Models for Music: A Survey

Yinghao Ma, Anders Øland, Anton Ragni +40 authors

A review of foundation models in music, including large language models and latent diffusion models, highlights their impact, architectural choices, and the need for ethical considerations in music applications.

44large language modelslatent diffusion modelsHF ↗arXiv ↗
228

TransformerFAM: Feedback attention is working memory

Dongseong Hwang, Weiran Wang, Zhuoyuan Huo +2 authors

Feedback Attention Memory (FAM) enhances Transformer architecture by enabling long-context processing without additional weights, significantly improving performance on large sequences across various model sizes.

43TransformersFeedback Attention MemoryHF ↗arXiv ↗
229

MindSearch: Mimicking Human Minds Elicits Deep AI Searcher

Zehui Chen, Kuikun Liu, Qiuchen Wang +4 authors

MindSearch, an LLM-based multi-agent framework, improves web information seeking and integration through parallel processing and hierarchical retrieval, achieving better performance than existing solutions.

43Large Language ModelsLLMsHF ↗arXiv ↗
230

In-Context Learning Creates Task Vectors

Roee Hendel, Mor Geva, Amir Globerson

In-Context Learning in Large Language Models can be understood as compressing a training set into a task vector that modulates a transformer for output generation.

43in-context learninglarge language modelsHF ↗arXiv ↗
231

Agents: An Open-source Framework for Autonomous Language Agents

Wangchunshu Zhou, Yuchen Eleanor Jiang, Long Li +14 authors

Agents is an open-source library that facilitates the creation, customization, and deployment of autonomous language agents with features like planning, memory, and tool usage, aiming to make advances in large language models accessible to a broader audience.

43large language modelsautonomous language agentsHF ↗arXiv ↗
232

How to Train Data-Efficient LLMs

Noveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang +6 authors

Data-efficient methods like Ask-LLM and Density sampling improve model quality and training efficiency in large language models by optimizing data selection and coverage.

43large language modelsdata-efficientHF ↗arXiv ↗
233

VILA^2: VILA Augmented VILA

Yunhao Fang, Ligeng Zhu, Yao Lu +6 authors

A novel data augmentation approach iteratively improves visual language model data quality and performance using self-augmentation and specialist-augmentation, leading to state-of-the-art results on MMMU tasks.

41visual language modelslarge language modelsHF ↗arXiv ↗
235

McEval: Massively Multilingual Code Evaluation

Linzheng Chai, Shukai Liu, Jian Yang +15 authors

A multilingual code benchmark covering 40 programming languages with 16K test samples is introduced to advance code language model research, along with a multilingual coder model and instruction corpora.

41large language modelscode understandingHF ↗arXiv ↗
238

sDPO: Don't Use Your Data All at Once

Dahyun Kim, Yungi Kim, Wonho Song +4 authors

A stepwise direct preference optimization approach improves the alignment of large language models with human preferences and enhances their performance.

41large language modelsLLMHF ↗arXiv ↗
240

HARE: HumAn pRiors, a key to small language model Efficiency

Lingyun Zhang, Bin jin, Gaojian Ge +7 authors

A principle for leveraging human priors in data construction is proposed to improve small language models, demonstrating favorable performance on large benchmarks in resource-constrained settings.

40large language modelssmall language modelsHF ↗arXiv ↗
12 / 16

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号