Qwen-Image Technical Report
Chenfei Wu, Jiahao Li, Jingren Zhou +36 authors
Qwen-Image, an image generation model, advances text rendering and image editing through a comprehensive data pipeline, progressive training, and dual-encoding mechanism.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23 authors
DINOv3, a self-supervised learning model, achieves superior performance across various vision tasks by scaling datasets and models, addressing dense feature degradation, and enhancing flexibility with post-hoc strategies.
50 篇论文 · 按点赞排序
Chenfei Wu, Jiahao Li, Jingren Zhou +36 authors
Qwen-Image, an image generation model, advances text rendering and image editing through a comprehensive data pipeline, progressive training, and dual-encoding mechanism.
Lei Bai, Zhongrui Cai, Maosong Cao +172 authors
Intern-S1, a multimodal Mixture-of-Experts model with extensive pre-training and reinforcement learning, achieves top-tier performance in general reasoning and outperforms closed-source models in scientific tasks.
Chengshuai Zhao, Zhen Tan, Pingchuan Ma +5 authors
CoT reasoning in LLMs is found to be limited by the distribution discrepancy between training and test data, suggesting it is not a robust form of reasoning.
Weiyun Wang, Zhangwei Gao, Lixin Gu +58 authors
InternVL 3.5 introduces Cascade RL, ViR, and DvD to enhance reasoning, efficiency, and performance in multimodal models.
GLM-4. 5 Team, Aohan Zeng, Xin Lv +168 authors
GLM-4.5, a Mixture-of-Experts large language model with 355B parameters, achieves strong performance across agentic, reasoning, and coding tasks using multi-stage training and reinforcement learning.
Yongliang Wu, Yizhou Zhou, Zhou Ziheng +7 authors
Dynamic Fine-Tuning (DFT) improves the generalization of Large Language Models (LLMs) by dynamically rescaling gradients, outperforming standard Supervised Fine-Tuning (SFT) and showing competitive results in offline reinforcement learning.
Zhiliang Peng, Jianwei Yu, Wenhui Wang +10 authors
VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.
Shunyu Liu, Minghao Liu, Huichi Zhou +29 authors
VeriGUI is a novel dataset for evaluating GUI agents in long-horizon tasks, emphasizing long-chain complexity and subtask-level verifiability.
Huichi Zhou, Yihang Chen, Siyuan Guo +8 authors
A novel memory-augmented reinforcement learning paradigm enables adaptive LLM agents to continually learn without fine-tuning, using episodic memory and a neural case-selection policy.
NextStep Team, Chunrui Han, Guopeng Li +47 authors
NextStep-1, a 14B autoregressive model with a 157M flow matching head, achieves state-of-the-art performance in text-to-image generation and image editing by processing discrete text tokens and continuous image tokens.
Runqi Qiao, Qiuna Tan, Peiqing Yang +11 authors
We-Math 2.0 enhances MLLMs' mathematical reasoning through a structured knowledge system, model-centric data space modeling, and reinforcement learning, demonstrating competitive performance on benchmarks.
Xinyu Geng, Peng Xia, Zhen Zhang +11 authors
WebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.
Xufang Luo, Yuge Zhang, Zhiyuan He +5 authors
Agent Lightning is a flexible RL framework for training LLMs in various agents, using a hierarchical RL algorithm and decoupling execution from training to handle complex interactions.
Yuxuan Song, Zheng Zhang, Cheng Luo +19 authors
Seed Diffusion Preview, a discrete-state diffusion language model, achieves fast inference speeds through parallel generation, outperforming Mercury and Gemini Diffusion in speed and quality.
Chengsong Huang, Wenhao Yu, Xiaoyang Wang +6 authors
R-Zero is a self-evolving framework that autonomously generates and learns from its own training data, improving reasoning capabilities in LLMs without human-curated tasks.
Weizhen Li, Jianbo Lin, Zhuosong Jiang +27 authors
Chain-of-Agents (CoA) paradigm enables end-to-end complex problem-solving in LLMs through dynamic agent activation, improving performance via multi-agent distillation and agentic reinforcement learning.
Ning Shang, Yifei Liu, Yi Zhu +12 authors
rStar2-Agent, a 14B math reasoning model trained with agentic reinforcement learning, achieves state-of-the-art performance by efficiently handling complex problem-solving with advanced cognitive behaviors and minimal computational resources.
Xiao Liang, Zhongzhi Li, Yeyun Gong +4 authors
An online self-play strategy with variational problem synthesis for RLVR training maintains policy entropy and improves Pass@k performance on reasoning benchmarks.
Wenhan Liu, Xinyu Ma, Weiwei Sun +4 authors
A reasoning-intensive reranker, ReasonRank, achieves state-of-the-art performance in passage ranking tasks by using synthesized training data and a two-stage post-training approach with reinforcement learning.
Shiyin Lu, Yang Li, Yu Xia +39 authors
Ovis2.5, a native-resolution vision transformer with multimodal reasoning, achieves state-of-the-art performance on various benchmarks through advanced training techniques and efficient scaling methods.
Luoxin Chen, Jinming Gu, Liankai Huang +33 authors
Seed-Prover, a lemma-style reasoning model using Lean, achieves high performance in formal theorem proving and automated mathematical reasoning through iterative refinement and specialized geometry support.
Ryan Wong, Jiawei Wang, Junjie Zhao +10 authors
WideSearch is a new benchmark evaluating the reliability of automated search agents in large-scale information collection tasks, revealing significant deficiencies in current systems.
Jinyuan Fang, Yanwen Peng, Xi Zhang +12 authors
A survey of self-evolving AI agents that adapt to dynamic environments through automatic enhancement based on interaction data and feedback.
Yuchen Fan, Kaiyan Zhang, Heng Zhou +15 authors
LLMs can serve as efficient simulators for RL tasks by leveraging internal knowledge, reducing reliance on external search engines and improving sim-to-real transfer.
Tianqing Fang, Zhisong Zhang, Xiaoyang Wang +10 authors
Cognitive Kernel-Pro is an open-source multi-module agent framework that enhances AI agent robustness and performance through data curation and novel test-time strategies, achieving state-of-the-art results.
Neta Glazer, Yael Segal-Feldman, Hilit Segev +6 authors
Interpretability methods like logit lens, linear probing, and activation patching are applied to ASR to uncover internal dynamics, repetition hallucinations, and semantic biases, enhancing model transparency and robustness.
Yichao Fu, Xuewei Wang, Yuandong Tian +1 authors
DeepConf enhances reasoning efficiency and performance by filtering low-quality reasoning traces using model-internal confidence signals, achieving high accuracy and reducing token generation.
Yibin Wang, Zhimin Li, Yuhang Zang +6 authors
Pref-GRPO, a pairwise preference reward-based GRPO method, enhances text-to-image generation by mitigating reward hacking and improving stability, while UniGenBench provides a comprehensive benchmark for evaluating T2I models.
Shuaijie She, Yu Bao, Yu Lu +7 authors
DuPO is a dual learning framework that generates annotation-free feedback using a generalized duality, enhancing performance across various tasks without relying on costly labels.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号