TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

multimodal understanding 相关论文

16 篇论文 · 按点赞排序

01

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model

Inclusion AI, Tiwei Bie, Haoxing Chen +15 authors

LLaDA2.0-Uni is a unified discrete diffusion language model that integrates multimodal understanding and generation through a semantic discrete tokenizer, MoE-based backbone, and diffusion decoder, achieving performance comparable to specialized vision-language models while enabling efficient inference and high-fidelity image generation.

245discrete diffusionlarge language modelHF ↗arXiv ↗
02

Seed1.5-VL Technical Report

Dong Guo, Faming Wu, Feida Zhu +194 authors

Seed1.5-VL, a vision-language foundation model combining a vision encoder and a large MoE LLM, achieves state-of-the-art performance across various benchmarks and excels in multimodal reasoning tasks such as visual puzzles.

157vision-language foundation modelvision encoderHF ↗arXiv ↗
03

Emerging Properties in Unified Multimodal Pretraining

Chaorui Deng, Deyao Zhu, Kunchang Li +9 authors

BAGEL, an open-source foundational model trained on diverse multimodal data, significantly outperforms existing models in both generation and understanding tasks.

136multimodal understandingmultimodal generationHF ↗arXiv ↗
05

Aria: An Open Multimodal Native Mixture-of-Experts Model

Dongxu Li, Yudong Liu, Haoning Wu +7 authors

Aria is an open multimodal native AI model with best-in-class performance across various tasks, designed with a mixture-of-experts architecture and pre-trained through a four-stage pipeline.

111mixture-of-expert modelvisual tokenHF ↗arXiv ↗
07

MMaDA: Multimodal Large Diffusion Language Models

Ling Yang, Ye Tian, Bowen Li +4 authors

MMaDA, a multimodal diffusion foundation model, achieves superior performance through a unified architecture, mixed long chain-of-thought fine-tuning, and a unified policy-gradient-based RL algorithm.

99multimodal diffusion foundation modelsunified diffusion architectureHF ↗arXiv ↗
10

Magistral

Mistral-AI, Abhinav Rastogi, Albert Q. Jiang +97 authors

Magistral, a scalable reinforcement learning pipeline, demonstrates that RL can enhance multimodal understanding and instruction following in large language models without requiring existing RL traces.

69reinforcement learningRLHF ↗arXiv ↗
12

Ovis-U1 Technical Report

Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9 authors

Ovis-U1, a 3-billion-parameter unified model, integrates multimodal understanding, text-to-image generation, and image editing using a diffusion-based visual decoder and bidirectional token refiner, achieving state-of-the-art performance across various benchmarks.

63diffusion-based visual decoderbidirectional token refinerHF ↗arXiv ↗

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号