TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

benchmarks 相关论文

23 篇论文 · 按点赞排序

01

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu +22 authors

Agentic reinforcement learning transforms large language models into autonomous decision-making agents by leveraging temporally extended POMDPs, enhancing capabilities like planning and reasoning through reinforcement learning.

239agentic reinforcement learningLLM RLHF ↗arXiv ↗
02

Agentic Reasoning for Large Language Models

Tianxin Wei, Ting-Wei Li, Zhining Liu +26 authors

Agentic reasoning redefines large language models as autonomous agents capable of planning, acting, and learning through continuous interaction in dynamic environments across single-agent and multi-agent frameworks.

207large language modelsagentic reasoningHF ↗arXiv ↗
06

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +661 authors

HLE is a challenging multi-modal benchmark that highlights the limitations of current LLMs in closed-ended academic questions.

78large language model (LLM)benchmarksHF ↗arXiv ↗
11

LLaMA Pro: Progressive LLaMA with Block Expansion

Chengyue Wu, Yukang Gan, Yixiao Ge +5 authors

A new post-pretraining method using expanded Transformer blocks for Large Language Models improves knowledge without catastrophic forgetting, yielding LLaMA Pro-8.3B that excels in general tasks, programming, and mathematics.

54Large Language ModelsLLMsHF ↗arXiv ↗
12

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

NVIDIA, Alisson Azzolini, Hannah Brandon +42 authors

Cosmos-Reason1 models, using hierarchical and two-dimensional ontologies for physical common sense and embodied reasoning, generate embodied decisions through multimodal large language models trained in vision and Physical AI stages.

52Physical AIreasoningHF ↗arXiv ↗
14

Evaluating and Aligning CodeLLMs on Human Preference

Jian Yang, Jiaxi Yang, Ke Jin +7 authors

A human-curated benchmark (CodeArena) and a large synthetic instruction corpus (SynCode-Instruct) are introduced to evaluate code LLMs based on human preference alignment, revealing performance differences between open-source and proprietary models.

48code large language modelscode generationHF ↗arXiv ↗
20

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97 authors

InternLM2 is an open-source LLM that outperforms predecessors through innovative pre-training and optimization techniques, including Supervised Fine-Tuning and Conditional Online Reinforcement Learning from Human Feedback.

35Large Language ModelsLLMsHF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号