TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

multimodal large language models 相关论文

70 篇论文 · 按点赞排序

41

Personalized Visual Instruction Tuning

Renjie Pi, Jianshu Zhang, Tianyang Han +3 authors

A new framework called Personalized Visual Instruction Tuning (PVIT) enhances multimodal large language models to recognize and engage with specific individuals in images, utilizing a curated dataset and benchmarks for evaluation.

70multimodal large language modelsMLLMsHF ↗arXiv ↗
44

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

Jiaqi Tang, Jianmin Chen, Wei Wei +7 authors

A novel framework, Robust-R1, enhances multimodal large language models' robustness to visual degradations through explicit modeling, supervised fine-tuning, reward-driven alignment, and dynamic reasoning depth scaling, achieving state-of-the-art performance on real-world degradation benchmarks.

68multimodal large language modelsvisual degradationsHF ↗arXiv ↗
48

Kosmos-2.5: A Multimodal Literate Model

Tengchao Lv, Yupan Huang, Jingye Chen +11 authors

Kosmos-2.5, a unified multimodal model, generates spatially-aware and structured text from text-intensive images using a Transformer architecture and task-specific prompts.

56multimodal literate modelmachine readingHF ↗arXiv ↗
49

Needle In A Multimodal Haystack

Weiyun Wang, Shuibo Zhang, Yiming Ren +13 authors

NN-NIAH is a benchmark that evaluates MLLMs' comprehension of long multimodal documents across retrieval, counting, and reasoning tasks.

55multimodal large language modelsMLLMsHF ↗arXiv ↗
52

MIO: A Foundation Model on Multimodal Tokens

Zekun Wang, King Zhu, Chunpu Xu +14 authors

MIO, a novel foundation model, achieves competitive performance in multimodal tasks through an end-to-end, autoregressive approach using causal multimodal modeling.

53MIOmultimodal tokensHF ↗arXiv ↗
53

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

NVIDIA, Alisson Azzolini, Hannah Brandon +42 authors

Cosmos-Reason1 models, using hierarchical and two-dimensional ontologies for physical common sense and embodied reasoning, generate embodied decisions through multimodal large language models trained in vision and Physical AI stages.

52Physical AIreasoningHF ↗arXiv ↗
56

MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

Run Luo, Haonan Zhang, Longze Chen +13 authors

MMEvol, a multimodal instruction data evolution framework, enhances the capabilities of Multimodal Large Language Models by generating diverse and complex image-text instruction datasets, leading to improved performance across various vision-language tasks.

49Multimodal Large Language ModelsMMEvolHF ↗arXiv ↗
60

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Yifan Shen, Jian Xu, Boyi Li +6 authors

ChronoVision improves visual temporal reasoning by aligning latent imagery with logic through reconstructive prediction, ROI attention, and reinforcement learning, achieving strong results on video reasoning benchmarks.

39multimodal large language modelsReconstructive Visual HeadHF ↗arXiv ↗
3 / 4

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号