TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

mllms 相关论文

29 篇论文 · 按点赞排序

04

Ovis2.5 Technical Report

Shiyin Lu, Yang Li, Yu Xia +39 authors

Ovis2.5, a native-resolution vision transformer with multimodal reasoning, achieves state-of-the-art performance on various benchmarks through advanced training techniques and efficient scaling methods.

116vision transformernative-resolutionHF ↗arXiv ↗
07

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Yuan Yao, Tianyu Yu, Ao Zhang +20 authors

MiniCPM-V presents a series of efficient Multimodal Large Language Models optimized for end-side deployment, offering high performance and practical usability compared to larger models.

96Multimodal Large Language ModelsMLLMsHF ↗arXiv ↗
08

Law of Vision Representation in MLLMs

Shijia Yang, Bohan Zhai, Quanzeng You +3 authors

Correlation between cross-modal alignment and vision representation improves performance in multimodal large language models, enabling identification and training of optimal vision representation with reduced computational cost.

95cross-modal alignmentvision representationHF ↗arXiv ↗
12

Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Keen You, Haotian Zhang, Eldon Schoop +5 authors

Ferret-UI, a multimodal large language model tailored for mobile UI screens, enhances understanding and interaction through region annotations and a comprehensive dataset of UI tasks, outperforming existing models including GPT-4V.

83multimodal large language modelsMLLMsHF ↗arXiv ↗
13

Thyme: Think Beyond Images

Yi-Fan Zhang, Xingyu Lu, Shukang Yin +17 authors

Thyme, a novel paradigm, enables MLLMs to autonomously perform image manipulations and computations, enhancing performance in perception and reasoning tasks through a two-stage training strategy and GRPO-ATS algorithm.

81MLLMsthink with imagesHF ↗arXiv ↗
14

Video-R1: Reinforcing Video Reasoning in MLLMs

Kaituo Feng, Kaixiong Gong, Bohao Li +5 authors

Video-R1, leveraging rule-based reinforcement learning and temporal information, enhances video reasoning in multimodal large language models using a combination of video and image data.

79rule-based reinforcement learningRLHF ↗arXiv ↗
16

Personalized Visual Instruction Tuning

Renjie Pi, Jianshu Zhang, Tianyang Han +3 authors

A new framework called Personalized Visual Instruction Tuning (PVIT) enhances multimodal large language models to recognize and engage with specific individuals in images, utilizing a curated dataset and benchmarks for evaluation.

70multimodal large language modelsMLLMsHF ↗arXiv ↗
20

Needle In A Multimodal Haystack

Weiyun Wang, Shuibo Zhang, Yiming Ren +13 authors

NN-NIAH is a benchmark that evaluates MLLMs' comprehension of long multimodal documents across retrieval, counting, and reasoning tasks.

55multimodal large language modelsMLLMsHF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号