TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

benchmark 相关论文

25 篇论文 · 按点赞排序

01

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +305 authors

Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on long-term, economically valuable real-world tasks across 13 industry clusters with 1K+ tasks, revealing significant gaps between benchmark performance and practical deployment.

385AI agentsbenchmarkHF ↗arXiv ↗
02

CLEAR: Character Unlearning in Textual and Visual Modalities

Alexey Dontsov, Dmitrii Korzh, Alexey Zhavoronkin +6 authors

CLEAR benchmark evaluates multimodal unlearning methods across textual and visual data, highlighting challenges and demonstrating the effectiveness of $\ell_1$ regularization on LoRA weights in mitigating catastrophic forgetting.

209Machine UnlearningMUHF ↗arXiv ↗
03

SAM 3D: 3Dfy Anything in Images

SAM 3D Team, Xingyu Chen, Fu-Jen Chu +20 authors

SAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.

117generative model3D object reconstructionHF ↗arXiv ↗
04

WideSearch: Benchmarking Agentic Broad Info-Seeking

Ryan Wong, Jiawei Wang, Junjie Zhao +10 authors

WideSearch is a new benchmark evaluating the reliability of automated search agents in large-scale information collection tasks, revealing significant deficiencies in current systems.

113Large Language Modelsautomated search agentsHF ↗arXiv ↗
10

Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Keen You, Haotian Zhang, Eldon Schoop +5 authors

Ferret-UI, a multimodal large language model tailored for mobile UI screens, enhances understanding and interaction through region annotations and a comprehensive dataset of UI tasks, outperforming existing models including GPT-4V.

83multimodal large language modelsMLLMsHF ↗arXiv ↗
11

Personalized Visual Instruction Tuning

Renjie Pi, Jianshu Zhang, Tianyang Han +3 authors

A new framework called Personalized Visual Instruction Tuning (PVIT) enhances multimodal large language models to recognize and engage with specific individuals in images, utilizing a curated dataset and benchmarks for evaluation.

70multimodal large language modelsMLLMsHF ↗arXiv ↗
12

TrustLLM: Trustworthiness in Large Language Models

Lichao Sun, Yue Huang, Haoran Wang +64 authors

This study assesses the trustworthiness of large language models across various dimensions, including truthfulness, safety, fairness, robustness, privacy, and machine ethics, finding a positive correlation with utility and highlighting differences between proprietary and open-source models.

69TrustLLMlarge language modelsHF ↗arXiv ↗
20

Weaver: Foundation Models for Creative Writing

Tiannan Wang, Jiamin Chen, Qingrui Jia +43 authors

Weaver, a family of specialized large language models, achieves superior writing capabilities through pre-training and fine-tuning methods, surpasses GPT-4 in various writing tasks, and supports retrieval-augmented generation and tool usage.

46large language modelspre-trainingHF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号