TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

text-to-image 相关论文

16 篇论文 · 按点赞排序

01

Qwen-Image Technical Report

Chenfei Wu, Jiahao Li, Jingren Zhou +36 authors

Qwen-Image, an image generation model, advances text rendering and image editing through a comprehensive data pipeline, progressive training, and dual-encoding mechanism.

276data pipelineprogressive trainingHF ↗arXiv ↗
02

From Editor to Dense Geometry Estimator

JiYuan Wang, Chunyu Lin, Lei Sun +6 authors

FE2E, a framework using a Diffusion Transformer for dense geometry prediction, outperforms generative models in zero-shot monocular depth and normal estimation with improved performance and efficiency.

96text-to-imagedense predictionHF ↗arXiv ↗
07

DanceOPD: On-Policy Generative Field Distillation

Wei Zhou, Xiongwei Zhu, Zelin Xu +8 authors

A novel on-policy generative field distillation framework called DanceOPD is proposed to unify text-to-image generation, local editing, and global editing capabilities in flow-matching models through capability-specific routing and velocity-based training.

81generative field distillationflow-matching modelsHF ↗arXiv ↗
08

RewardDance: Reward Scaling in Visual Generation

Jie Wu, Yu Gao, Zilyu Ye +9 authors

RewardDance is a scalable reward modeling framework that aligns with VLM architectures, enabling effective scaling of RMs and resolving reward hacking issues in generation models.

73CLIP-based RMsBradley-Terry lossesHF ↗arXiv ↗
10

An Empirical Study of GPT-4o Image Generation Capabilities

Sixiang Chen, Jinbin Bai, Zhuoran Zhao +16 authors

An empirical study of GPT-4o's image generation capabilities across multiple tasks reveals its strengths and limitations compared to other models, highlighting the importance of architectural design and data scaling in unified generative frameworks.

64GANdiffusion modelsHF ↗arXiv ↗
11

MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation

Weimin Wang, Jiawei Liu, Zhijie Lin +9 authors

MagicVideo-V2 generates high-fidelity and smooth videos from text using an integrated pipeline that includes text-to-image, video motion generation, and frame interpolation modules, outperforming existing systems in user evaluations.

49text-to-imagevideo motion generatorHF ↗arXiv ↗
12

Kandinsky 3.0 Technical Report

Vladimir Arkhipkin, Andrei Filatov, Viacheslav Vasilev +6 authors

Kandinsky 3.0, a large-scale text-to-image model based on latent diffusion, improves quality and realism through a larger architecture and advanced text understanding.

45latent diffusionU-NetHF ↗arXiv ↗
14

ChatAnything: Facetime Chat with LLM-Enhanced Personas

Yilin Zhao, Xinbin Yuan, Shanghua Gao +4 authors

A framework for generating anthropomorphized personas with diverse voices and appearances from text descriptions using LLMs and generative models, with improved face landmark detection for automatic animation.

35LLM-based charactersin-context learningHF ↗arXiv ↗

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
11 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
13
37 篇论文
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号