TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

vision-language models 相关论文

54 篇论文 · 按点赞排序

41

DataComp-VLM: Improved Open Datasets for Vision-Language Models

Matteo Farina, Vishaal Udandarao, Thao Nguyen +33 authors

DataComp for VLMs (DCVLM) establishes a comprehensive benchmark for evaluating data curation strategies in vision-language models, demonstrating that data mixing rather than filtering significantly improves model performance at scale.

52Vision-Language Modelsdata curationHF ↗arXiv ↗
44

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Qifeng Zhang, Kaixiang Huang, Heng Dong +6 authors

The study introduces a video benchmark requiring global spatial reasoning across long videos and reveals that vision-language models struggle to build consistent global scene representations despite strong local perception.

46global spatial intelligenceVQA benchmarkHF ↗arXiv ↗
45

Articulated Object Reconstruction from Rest-State Observation

Daeun Lee, Jaeah Lee, Woosung Kim +2 authors

A rest-state framework reconstructs articulated objects from a single closed configuration by fusing vision-language outputs into consistent part meshes and validating synthesized motion hypotheses via geometric consistency.

46digital twinsarticulated object reconstructionHF ↗arXiv ↗
48

3D-LLM: Injecting the 3D World into Large Language Models

Yining Hong, Haoyu Zhen, Peihao Chen +4 authors

A new family of 3D-LLMs is introduced to perform 3D-related tasks by leveraging 3D point clouds and features, outperforming state-of-the-art baselines in tasks such as 3D question answering and captioning.

40LLMsVision-Language ModelsHF ↗arXiv ↗
51

HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models

Fuxiao Liu, Tianrui Guan, Zongxia Li +4 authors

HallusionBench is a benchmark that highlights language hallucination and visual illusion issues in vision-language models (VLMs), showcasing the limitations of current state-of-the-art models like GPT-4V and LLaVA-1.5.

27Large language modelsvision modelsHF ↗arXiv ↗
52

Xiaomi-GUI-0 Technical Report

Wanxia Cao, Chengzhen Duan, Pei Fu +27 authors

A native multimodal GUI agent trained in real-device environments demonstrates superior performance and stability compared to traditional benchmark-based approaches.

19vision-language modelsinterface actionsHF ↗arXiv ↗
3 / 3

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号