TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

multimodal large language model 相关论文

17 篇论文 · 按点赞排序

03

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Zhe Chen, Weiyun Wang, Yue Cao +37 authors

InternVL 2.5, an advanced multimodal large language model, showcases competitive performance across various benchmarks, including multimodal reasoning and understanding, and is the first open-source model to surpass 70% on the MMMU benchmark using Chain-of-Thought reasoning.

162multimodal large language modelvision encodersHF ↗arXiv ↗
05

Web-Shepherd: Advancing PRMs for Reinforcing Web Agents

Hyungjoo Chae, Sunghwan Kim, Junhee Cho +18 authors

The paper introduces Web-Shepherd, a process reward model for web navigation, which improves accuracy and cost-effectiveness in step-level trajectory assessment compared to existing multimodal large language models.

105multimodal large language modelprocess reward modelHF ↗arXiv ↗
06

Baichuan-Omni Technical Report

Yadong Li, Haoze Sun, Mingan Lin +24 authors

Baichuan-Omni, a 7B open-source Multimodal Large Language Model, excels in processing image, video, audio, and text, showcasing competitive performance across multimodal benchmarks.

88Multimodal Large Language Modelmultimodal training schemaHF ↗arXiv ↗
09

UniVideo: Unified Understanding, Generation, and Editing for Videos

Cong Wei, Quande Liu, Zixuan Ye +5 authors

UniVideo, a dual-stream framework combining a Multimodal Large Language Model and a Multimodal DiT, extends unified modeling to video generation and editing, achieving state-of-the-art performance and supporting task composition and generalization.

81Multimodal Large Language ModelMultimodal DiTHF ↗arXiv ↗
13

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Qian Chen, Yafeng Chen, Yanni Chen +33 authors

MinMo, a multimodal large language model, integrates speech and text processing to achieve state-of-the-art performance in voice comprehension and generation, while enabling full-duplex conversation and instruction-following capabilities.

54multimodal large language modelnative modelsHF ↗arXiv ↗
14

VITA: Towards Open-Source Interactive Omni Multimodal LLM

Chaoyou Fu, Haojia Lin, Zuwei Long +12 authors

VITA, an open-source Multimodal Large Language Model, excels in processing Video, Image, Text, and Audio with seamless interaction, showcasing advancements in multimodal understanding and human-computer interaction.

50Multimodal Large Language ModelMixtralHF ↗arXiv ↗

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号