Qwen-Image Technical Report
Chenfei Wu, Jiahao Li, Jingren Zhou +36 authors
Qwen-Image, an image generation model, advances text rendering and image editing through a comprehensive data pipeline, progressive training, and dual-encoding mechanism.
Trends · 研究趋势
数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。
Chenfei Wu, Jiahao Li, Jingren Zhou +36 authors
Qwen-Image, an image generation model, advances text rendering and image editing through a comprehensive data pipeline, progressive training, and dual-encoding mechanism.
JiYuan Wang, Chunyu Lin, Lei Sun +6 authors
FE2E, a framework using a Diffusion Transformer for dense geometry prediction, outperforms generative models in zero-shot monocular depth and normal estimation with improved performance and efficiency.
Yibin Wang, Zhimin Li, Yuhang Zang +6 authors
Pref-GRPO, a pairwise preference reward-based GRPO method, enhances text-to-image generation by mitigating reward hacking and improving stability, while UniGenBench provides a comprehensive benchmark for evaluating T2I models.
Kwanyoung Kim, Byeongsu Sim
PLADIS leverages sparse attention in cross-attention layers to enhance pre-trained text-to-image diffusion models, improving text alignment and human preference without additional training.
Hengyuan Xu, Wei Cheng, Peng Xing +8 authors
A diffusion-based model addresses copy-paste artifacts in text-to-image generation by using a large-scale paired dataset and a contrastive identity loss to balance identity fidelity and variation.
Valerii Startsev, Alexander Ustyuzhanin, Alexey Kirillov +2 authors
A new method using a pre-trained generative model helps construct a high-impact SFT dataset, Alchemist, which improves the generative quality of text-to-image models while maintaining diversity.
Wei Zhou, Xiongwei Zhu, Zelin Xu +8 authors
A novel on-policy generative field distillation framework called DanceOPD is proposed to unify text-to-image generation, local editing, and global editing capabilities in flow-matching models through capability-specific routing and velocity-based training.
Jie Wu, Yu Gao, Zilyu Ye +9 authors
RewardDance is a scalable reward modeling framework that aligns with VLM architectures, enabling effective scaling of RMs and resolving reward hacking issues in generation models.
Junying Chen, Zhenyang Cai, Pengcheng Chen +5 authors
ShareGPT-4o-Image and Janus-4o enable open research in photorealistic, instruction-aligned image generation through a large dataset and multimodal model.
Sixiang Chen, Jinbin Bai, Zhuoran Zhao +16 authors
An empirical study of GPT-4o's image generation capabilities across multiple tasks reveals its strengths and limitations compared to other models, highlighting the importance of architectural design and data scaling in unified generative frameworks.
Weimin Wang, Jiawei Liu, Zhijie Lin +9 authors
MagicVideo-V2 generates high-fidelity and smooth videos from text using an integrated pipeline that includes text-to-image, video motion generation, and frame interpolation modules, outperforming existing systems in user evaluations.
Vladimir Arkhipkin, Andrei Filatov, Viacheslav Vasilev +6 authors
Kandinsky 3.0, a large-scale text-to-image model based on latent diffusion, improves quality and realism through a larger architecture and advanced text understanding.
Yaohui Wang, Xinyuan Chen, Xin Ma +17 authors
LaVie, a text-to-video framework integrating cascaded video latent diffusion models, achieves state-of-the-art video generation by leveraging temporal self-attentions and joint image-video fine-tuning.
Yilin Zhao, Xinbin Yuan, Shanghua Gao +4 authors
A framework for generating anthropomorphized personas with diverse voices and appearances from text descriptions using LLMs and generative models, with improved face landmark detection for automatic animation.
Chong Mou, Xintao Wang, Jiechong Song +2 authors
DragonDiffusion, a novel image editing method, enables precise manipulation of generated or real images using classifier guidance and multi-scale attention mechanisms without requiring fine-tuning.
Lingmin Ran, Xiaodong Cun, JiaWei Liu +5 authors
X-Adapter upgrades text-to-image diffusion models like SDXL to work with existing plugins without retraining, through feature remapping and a null-text training strategy.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号