TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

mmlu 相关论文

16 篇论文 · 按点赞排序

04

Qwen2 Technical Report

An Yang, Baosong Yang, Binyuan Hui +55 authors

The Qwen2 series, comprising 0.5 to 72 billion parameter models, surpasses prior open models across language understanding, generation, multilingualism, coding, math, and reasoning, with exceptional performance in benchmarks like MMLU, GPQA, HumanEval, GSM8K, BBH, MT-Bench, Arena-Hard, and LiveCodeBench.

175Mixture-of-Expertslanguage modelsHF ↗arXiv ↗
05

Qwen2.5-Omni Technical Report

Jin Xu, Zhifang Guo, Jinzheng He +11 authors

Qwen2.5-Omni is a multimodal model that processes text, images, audio, and video in a streaming fashion and generates text and speech using a dual-track architecture, achieving state-of-the-art performance on multimodal benchmarks.

173block-wise processingTMRoPE (Time-aligned Multimodal RoPE)HF ↗arXiv ↗
06

Diffusion Language Models are Super Data Learners

Jinjie Ni, Qian Liu, Longxu Dou +5 authors

Diffusion language models outperform autoregressive models in low-data settings due to any-order modeling, iterative bidirectional denoising, and Monte Carlo augmentation, and maintain advantages even at scale.

132diffusion language modelsautoregressive modelsHF ↗arXiv ↗
08

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +661 authors

HLE is a challenging multi-modal benchmark that highlights the limitations of current LLMs in closed-ended academic questions.

78large language model (LLM)benchmarksHF ↗arXiv ↗
09

Yi: Open Foundation Models by 01.AI

01. AI, Alex Young, Bei Chen +28 authors

The Yi model family, based on transformer architecture, showcases strong performance across benchmarks and modalities through optimized data and scalable infrastructure.

66language modelsmultimodal modelsHF ↗arXiv ↗
11

Make Your LLM Fully Utilize the Context

Shengnan An, Zexiong Ma, Zeqi Lin +2 authors

FILM-7B enhances long-context processing by using information-intensive training with synthesized datasets, improving retrieval and performance on real-world tasks without compromising short-context performance.

55information-intensive traininglong-context trainingHF ↗arXiv ↗
14

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +939 authors

Gemini, a family of multimodal models, achieves state-of-the-art performance across various benchmarks, including human-expert performance on MMLU, through advanced cross-modal reasoning and language understanding.

51multimodal modelscross-modal reasoningHF ↗arXiv ↗

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号