TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

multimodal large language models 相关论文

70 篇论文 · 按点赞排序

08

Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR

Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed +4 authors

Baseer, a vision-language model fine-tuned for Arabic document OCR, achieves state-of-the-art performance using a decoder-only strategy and a large-scale dataset, outperforming existing solutions with a WER of 0.25.

134Multimodal Large Language Modelsvision-language modelHF ↗arXiv ↗
09

Step-GUI Technical Report

Haolong Yan, Jia Wang, Xin Huang +94 authors

A self-evolving training pipeline with the Calibrated Step Reward System and GUI-MCP protocol improve GUI automation efficiency, accuracy, and privacy in real-world scenarios.

134multimodal large language modelsGUI automationHF ↗arXiv ↗
10

Kwai Keye-VL Technical Report

Kwai Keye Team, Biao Yang, Bin Wen +57 authors

Kwai Keye-VL, an 8-billion-parameter multimodal model, excels in short-video understanding and general vision-language tasks through a comprehensive pre-training and post-training process, including a five-mode data mixture and reinforcement learning.

133Multimodal Large Language ModelsKwai Keye-VLHF ↗arXiv ↗
17

MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization

Xiangyu Zhao, Junming Lin, Tianhao Liang +11 authors

Existing Multimodal Large Language Models show performance deficits in long-chain reflective reasoning, which is addressed by developing MM-HELIX-100K and Adaptive Hybrid Policy Optimization, leading to improved accuracy and generalization.

110Multimodal Large Language Modelslong-chain reflective reasoningHF ↗arXiv ↗
20

FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios

Xiangru Jian, Hao Xu, Wei Pang +13 authors

FORGE introduces a high-quality multimodal manufacturing dataset with fine-grained domain semantics to evaluate MLLMs on real-world tasks, revealing that domain-specific knowledge rather than visual grounding limits performance, and demonstrating that supervised fine-tuning on structured annotations significantly improves accuracy.

97Multimodal Large Language Modelsvisual groundingHF ↗arXiv ↗
1 / 4

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号