Qwen-Image Technical Report
Chenfei Wu, Jiahao Li, Jingren Zhou +36 authors
Qwen-Image, an image generation model, advances text rendering and image editing through a comprehensive data pipeline, progressive training, and dual-encoding mechanism.
Trends · 研究趋势
数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。
Chenfei Wu, Jiahao Li, Jingren Zhou +36 authors
Qwen-Image, an image generation model, advances text rendering and image editing through a comprehensive data pipeline, progressive training, and dual-encoding mechanism.
Yu Zhang, Ruiqi Li, Changhao Pan +3 authors
SwanTale is a multi-speaker expressive speech and audio generation model that supports both instruction-based and zero-shot synthesis with high-quality multi-modality outputs.
Renjie Pi, Grace Lam, Mohammad Shoeybi +3 authors
Researchers developed a synthetic task generation pipeline and analyzed data strategies to improve terminal agent performance, creating a large-scale dataset and models that outperform larger counterparts on benchmark tests.
Xin Xu, Clive Bai, Kai Yang +7 authors
Composition-RL improves reasoning capabilities by automatically composing multiple problems into new verifiable questions for reinforcement learning training.
Zhaorun Chen, Zhuokai Zhao, Kai Zhang +15 authors
DreamGym is a unified framework that synthesizes diverse experiences for scalable online RL training, improving agent performance and reducing real-world interactions.
Wei Shen, Jiangbo Pei, Yi Peng +7 authors
Skywork-R1V3, an open-source vision-language model, enhances visual reasoning through a post-training reinforcement learning framework, achieving state-of-the-art performance on multimodal reasoning tasks.
Omkar Thawakar, Dinura Dissanayake, Ketan More +12 authors
A framework for evaluating and improving step-by-step visual reasoning in large language models using a specialized benchmark and a novel multimodal model trained with curriculum learning.
Alexander Amini, Anna Banaszak, Harold Benoit +30 authors
LFM2, a family of compact foundation models, achieves high efficiency and performance on-device through hardware-in-the-loop architecture search and advanced training techniques, supporting various tasks including multimodal applications.
Kejian Zhu, Zhuoran Jin, Dongqi Huang +4 authors
Effective multimodal agent training is improved by selecting diverse environments via ability-aware selection and structuring difficulty through hierarchical curriculum learning.
Xiangpeng Wei, Haoran Wei, Huan Lin +15 authors
PolyLM, a multilingual LLM trained on 640 billion tokens, enhances multilingual capabilities through bilingual data and curriculum learning, outperforming other models on multilingual tasks while maintaining English performance.
Yecheng Jason Ma, William Liang, Guanzhi Wang +6 authors
Eureka, an LLM-powered reward design algorithm, generates human-level reward functions enhancing reinforcement learning for complex manipulation tasks and RLHF, including pen spinning with a simulated Shadow Hand.
Xincheng Wei, Yifan Ding, Yoshua Li +5 authors
DiagEvo improves language-model self-evolution by deriving training direction from internal failure history via hierarchical error-cause memory and double-confidence filtering, outperforming external-resource baselines.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号