TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

vision-language models 相关论文

54 篇论文 · 按点赞排序

21

UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

Zimo Wen, Boxiu Li, Wanbo Zhang +11 authors

Unified multimodal models show mixed performance in generation-to-understanding tasks, with specific subtasks benefiting from enhanced spatial and reasoning capabilities while overall performance lags behind specialized vision-language models.

88Unified multimodal modelsVision-Language ModelsHF ↗arXiv ↗
24

Self-Rewarding Vision-Language Model via Reasoning Decomposition

Zongxia Li, Wenhao Yu, Chengsong Huang +8 authors

Vision-SR1 uses reinforcement learning to enhance visual reasoning in vision-language models by decomposing the process into visual perception and language reasoning stages, improving accuracy and reducing hallucinations.

85vision-language modelsvisual hallucinationsHF ↗arXiv ↗
26

MiMo-VL Technical Report

Xiaomi LLM-Core Team, Zihao Yue, Zhenru Lin +71 authors

MiMo-VL-7B-SFT and MiMo-VL-7B-RL provide state-of-the-art general visual understanding and multimodal reasoning through four-stage pre-training and Mixed On-policy Reinforcement Learning, outperforming models with up to 78B parameters.

81vision-language modelsmultimodal reasoningHF ↗arXiv ↗
27

FineVision: Open Data Is All You Need

Luis Wiedmann, Orr Zohar, Amir Mahla +6 authors

FineVision, a large-scale and curated dataset, enhances vision-language models through rigorous data collection, de-duplication, and human oversight, leading to improved performance.

81vision-language modelsFineVisionHF ↗arXiv ↗
31

DeepSeek-OCR 2: Visual Causal Flow

Haoran Wei, Yaofeng Sun, Yukun Li

DeepSeek-OCR 2 introduces DeepEncoder V2 that dynamically reorders visual tokens based on semantic content, enabling more human-like causal reasoning in 2D image understanding through cascaded 1D causal structures.

72encoder-DeepEncoder V2visual tokensHF ↗arXiv ↗
33

Pixtral 12B

Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna +34 authors

Pixtral-12B, a 12-billion-parameter multimodal language model, excels in both natural language and image understanding, surpassing larger models and introducing an open-source benchmark for evaluation.

69multimodal language modelvision encoderHF ↗arXiv ↗
34

Yi: Open Foundation Models by 01.AI

01. AI, Alex Young, Bei Chen +28 authors

The Yi model family, based on transformer architecture, showcases strong performance across benchmarks and modalities through optimized data and scalable infrastructure.

66language modelsmultimodal modelsHF ↗arXiv ↗
35

One RL to See Them All: Visual Triple Unified Reinforcement Learning

Yan Ma, Linge Du, Xuyang Shen +7 authors

A unified reinforcement learning system, V-Triune, combines visual reasoning and perception tasks in vision-language models through a single training pipeline, achieving significant improvements across various tasks.

63visual triple unified reinforcement learningsample-level data formattingHF ↗arXiv ↗
36

Magma: A Foundation Model for Multimodal AI Agents

Jianwei Yang, Reuben Tan, Qianhui Wu +10 authors

Magma is a multimodal foundation model with both verbal intelligence and spatial-temporal intelligence, trained on diverse datasets to perform agentic tasks like UI navigation and robotic manipulation, outperforming specialized models.

58vision-language modelsspatial-temporal intelligenceHF ↗arXiv ↗
40

Unveiling Encoder-Free Vision-Language Models

Haiwen Diao, Yufeng Cui, Xiaotong Li +3 authors

EVE is an encoder-free vision-language model that achieves competitive performance on multiple benchmarks using a unified decoder and extra supervision.

55vision-language modelsVLMsHF ↗arXiv ↗
2 / 3

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号