TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

mixture-of-experts 相关论文

38 篇论文 · 按点赞排序

01

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +399 authors

Kimi K3 is a large-scale mixture-of-experts model with native vision and long-context capabilities that improves scaling efficiency and achieves strong performance across coding, reasoning, and agentic tasks.

509Mixture-of-ExpertsKimi Delta AttentionHF ↗arXiv ↗
02

Qwen2.5 Technical Report

Qwen, An Yang, Baosong Yang +39 authors

Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.

380large language modelspre-trainingHF ↗arXiv ↗
03

Group Sequence Policy Optimization

Chujie Zheng, Shixuan Liu, Mingze Li +9 authors

Group Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.

320Group Sequence Policy OptimizationGSPOHF ↗arXiv ↗
04

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +435 authors

ERNIE 5.0 is a production-scale trillion-parameter autoregressive model that unifies multimodal understanding and generation through sparse MoE architecture and elastic training.

269autoregressive foundation modelunified multimodal understandingHF ↗arXiv ↗
09

Kwai Keye-VL-2.0 Technical Report

Kwai Keye Team, Bin Wen, Changyi Liu +50 authors

Kwai Keye-VL-2.0-30B-A3B is an open-source Mixture-of-Experts multimodal foundation model that enables long-video understanding and agentic intelligence through DeepSeek Sparse Attention and specialized training infrastructure.

194Mixture-of-Expertsmultimodal foundation modelHF ↗arXiv ↗
10

LongCat-Flash-Thinking-2601 Technical Report

Meituan LongCat Team, Anchun Gui, Bei Li +159 authors

A 560-billion-parameter Mixture-of-Experts reasoning model achieves state-of-the-art performance on agentic benchmarks through a unified training framework combining domain-parallel expert training with fusion, along with enhancements for real-world robustness and complex reasoning.

183Mixture-of-Expertsagentic reasoningHF ↗arXiv ↗
11

Qwen2 Technical Report

An Yang, Baosong Yang, Binyuan Hui +55 authors

The Qwen2 series, comprising 0.5 to 72 billion parameter models, surpasses prior open models across language understanding, generation, multilingualism, coding, math, and reasoning, with exceptional performance in benchmarks like MMLU, GPQA, HumanEval, GSM8K, BBH, MT-Bench, Arena-Hard, and LiveCodeBench.

175Mixture-of-Expertslanguage modelsHF ↗arXiv ↗
12

Qwen3-VL Technical Report

Shuai Bai, Yuxuan Cai, Ruizhe Chen +61 authors

Qwen3-VL, a vision-language model, excels in text and multimodal understanding through advanced architectures and larger contexts, achieving superior performance across benchmarks.

164vision-language modelinterleaved contextsHF ↗arXiv ↗
14

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Shuqi Lu, Chaofan Li, Kun Luo +21 authors

AREX is a recursively self-improving deep research agent that verifies answers constraint-wise, compresses verified evidence into a compact state, and uses targeted follow-up research to refine results over long horizons.

154recursively self-improving agentsconstraint-wise verificationHF ↗arXiv ↗
16

Stabilizing Reinforcement Learning with LLMs: Formulation and Practices

Chujie Zheng, Kai Dang, Bowen Yu +7 authors

The paper provides a theoretical foundation for optimizing sequence-level rewards in reinforcement learning using token-level objectives, highlighting the importance of techniques like importance sampling correction, clipping, and Routing Replay for stabilizing training, especially with large language models.

107reinforcement learninglarge language modelsHF ↗arXiv ↗
19

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Lei Bai, Zongsheng Cao, Yang Chen +47 authors

Agents-A1, a 35B Mixture-of-Experts Agentic Model, achieves trillion-parameter-level performance through long-horizon trajectory scaling and heterogeneous agent ability scaling via a three-stage training approach involving supervised fine-tuning, domain-level teacher models, and multi-teacher distillation.

103Mixture-of-Expertsagentic modelHF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号