TensorX
返回文献探索

Paper · arXiv 2503.17352

OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement

Yihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng, Wei Wang, Kai-Wei Chang

24 upvotesMarch 21, 2025arXiv 预印本
AI 摘要

Iterative supervised fine-tuning and reinforcement learning improve reasoning capabilities in large vision-language models, enhancing performance on multimodal reasoning tasks like MathVista, MathVerse, and MathVision.

DeepSeek-R1reinforcement learningverifiable rewardslarge language modelslarge vision-language modelssupervised fine-tuningreasoning stepshigh-quality captionsvisual datasetsMathVistaMathVerseMathVision

Abstract

Recent advancements demonstrated by DeepSeek-R1 have shown that complex reasoning abilities in large language models (LLMs), including sophisticated behaviors such as self-verification and self-correction, can be achieved by RL with verifiable rewards and significantly improves model performance on challenging tasks such as AIME. Motivated by these findings, our study investigates whether similar reasoning capabilities can be successfully integrated into large vision-language models (LVLMs) and assesses their impact on challenging multimodal reasoning tasks. We consider an approach that iteratively leverages supervised fine-tuning (SFT) on lightweight training data and Reinforcement Learning (RL) to further improve model generalization. Initially, reasoning capabilities were distilled from pure-text R1 models by generating reasoning steps using high-quality captions of the images sourced from diverse visual datasets. Subsequently, iterative RL training further enhance reasoning skills, with each iteration's RL-improved model generating refined SFT datasets for the next round. This iterative process yielded OpenVLThinker, a LVLM exhibiting consistently improved reasoning performance on challenging benchmarks such as MathVista, MathVerse, and MathVision, demonstrating the potential of our strategy for robust vision-language reasoning. The code, model and data are held at https://github.com/yihedeng9/OpenVLThinker.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement | TensorX