TensorX
返回文献探索

Paper · arXiv 2512.08511

Thinking with Images via Self-Calling Agent

Wenxi Yang, Yuzhong Zhao, Fang Wan, Qixiang Ye

23 upvotesDecember 9, 2025arXiv 预印本
AI 摘要

sCoT, a language-only CoT paradigm with self-calling subagents, enhances visual reasoning performance and efficiency through group-relative policy optimization.

Chain-of-ThoughtCoTinterleaved multimodal CoTiMCoTreinforcement learningvisual reasoningparameter-sharing subagentsgroup-relative policy optimizationHR-Bench 4K

Abstract

Thinking-with-images paradigms have showcased remarkable visual reasoning capability by integrating visual information as dynamic elements into the Chain-of-Thought (CoT). However, optimizing interleaved multimodal CoT (iMCoT) through reinforcement learning remains challenging, as it relies on scarce high-quality reasoning data. In this study, we propose Self-Calling Chain-of-Thought (sCoT), a novel visual reasoning paradigm that reformulates iMCoT as a language-only CoT with self-calling. Specifically, a main agent decomposes the complex visual reasoning task to atomic subtasks and invokes its virtual replicas, i.e. parameter-sharing subagents, to solve them in isolated context. sCoT enjoys substantial training effectiveness and efficiency, as it requires no explicit interleaving between modalities. sCoT employs group-relative policy optimization to reinforce effective reasoning behavior to enhance optimization. Experiments on HR-Bench 4K show that sCoT improves the overall reasoning performance by up to 1.9% with sim 75% fewer GPU hours compared to strong baseline approaches. Code is available at https://github.com/YWenxi/think-with-images-through-self-calling.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号