TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

multimodal models 相关论文

28 篇论文 · 按点赞排序

01

Qwen2.5 Technical Report

Qwen, An Yang, Baosong Yang +39 authors

Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.

380large language modelspre-trainingHF ↗arXiv ↗
05

Qwen2 Technical Report

An Yang, Baosong Yang, Binyuan Hui +55 authors

The Qwen2 series, comprising 0.5 to 72 billion parameter models, surpasses prior open models across language understanding, generation, multilingualism, coding, math, and reasoning, with exceptional performance in benchmarks like MMLU, GPQA, HumanEval, GSM8K, BBH, MT-Bench, Arena-Hard, and LiveCodeBench.

175Mixture-of-Expertslanguage modelsHF ↗arXiv ↗
09

Emu3: Next-Token Prediction is All You Need

Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22 authors

Emu3, a transformer-based multimodal model trained exclusively with next-token prediction, outperforms existing diffusion and compositional models in generation and perception tasks.

99next-token predictionmultimodal modelsHF ↗arXiv ↗
10

V-Thinker: Interactive Thinking with Images

Runqi Qiao, Qiuna Tan, Minghan Yang +10 authors

V-Thinker, a multimodal reasoning assistant using reinforcement learning, enhances image-interactive thinking by synthesizing datasets and aligning perception for improved performance in vision-centric tasks.

98multimodal modelsimage interactionHF ↗arXiv ↗
14

Yi: Open Foundation Models by 01.AI

01. AI, Alex Young, Bei Chen +28 authors

The Yi model family, based on transformer architecture, showcases strong performance across benchmarks and modalities through optimized data and scalable infrastructure.

66language modelsmultimodal modelsHF ↗arXiv ↗
15

LLaVA-OneVision: Easy Visual Task Transfer

Bo Li, Yuanhan Zhang, Dong Guo +7 authors

LLaVA-OneVision is a unified multimodal model that advances performance across single-image, multi-image, and video scenarios with strong transfer learning capabilities.

61multimodal modelsLMMsHF ↗arXiv ↗
16

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +213 authors

Gemma 3 introduces vision capabilities, broader language coverage, and extended context length, featuring an optimized architecture and post-training enhancements to outperform previous versions.

58multimodal modelsvision understandingHF ↗arXiv ↗
20

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +939 authors

Gemini, a family of multimodal models, achieves state-of-the-art performance across various benchmarks, including human-expert performance on MMLU, through advanced cross-modal reasoning and language understanding.

51multimodal modelscross-modal reasoningHF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号