TensorX
返回文献探索

Paper · arXiv 2507.06167

Skywork-R1V3 Technical Report

Wei Shen, Jiangbo Pei, Yi Peng, Xuchen Song, Yang Liu, Jian Peng, Haofeng Sun, Yunzhuo Hao, Peiyu Wang, Yahui Zhou

75 upvotesJuly 8, 2025arXiv 预印本
AI 摘要

Skywork-R1V3, an open-source vision-language model, enhances visual reasoning through a post-training reinforcement learning framework, achieving state-of-the-art performance on multimodal reasoning tasks.

vision-language modelvisual reasoningLarge Language Modelspost-training RL frameworkconnector modulecross-modal alignmentmultimodal reasoning modelsentropy of critical reasoning tokensreinforcement finetuningcurriculum learning

Abstract

We introduce Skywork-R1V3, an advanced, open-source vision-language model (VLM) that pioneers a new approach to visual reasoning. Its key innovation lies in effectively transferring reasoning skills from text-only Large Language Models (LLMs) to visual tasks. The strong performance of Skywork-R1V3 primarily stems from our elaborate post-training RL framework, which effectively activates and enhances the model's reasoning ability, without the need for additional continue pre-training. Through this framework, we further uncover the fundamental role of the connector module in achieving robust cross-modal alignment for multimodal reasoning models. In addition, we introduce a unique indicator of reasoning capability, the entropy of critical reasoning tokens, which has proven highly effective for checkpoint selection during RL training. Skywork-R1V3 achieves state-of-the-art results on MMMU, significantly improving from 64.3% to 76.0%. This performance matches entry-level human capabilities. Remarkably, our RL-powered post-training approach enables even the 38B parameter model to rival top closed-source VLMs. The implementation successfully transfers mathematical reasoning to other subject-related reasoning tasks. We also include an analysis of curriculum learning and reinforcement finetuning strategies, along with a broader discussion on multimodal reasoning. Skywork-R1V3 represents a significant leap in multimodal reasoning, showcasing RL as a powerful engine for advancing open-source VLM capabilities.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Skywork-R1V3 Technical Report | TensorX