TensorX
返回文献探索

Paper · arXiv 2507.01951

Test-Time Scaling with Reflective Generative Model

Zixiao Wang, Yuxin Wang, Xiaorui Wang, Mengting Xing, Jie Gao, Jianjun Xu, Guangcan Liu, Chenhui Jin, Zhuo Wang, Shengzhuo Zhang, Hongtao Xie

108 upvotesJuly 2, 2025arXiv 预印本
AI 摘要

MetaStone-S1, a reflective generative model using a self-supervised process reward model, achieves high performance with reduced parameters and supports test time scaling.

reflective generative modelself-supervised process reward modelbackbone networktask-specific headspolicy modelprocess reward modeltest time scalingreasoning effort modesscaling lawtotal thinking computationTTS performance

Abstract

We introduce our first reflective generative model MetaStone-S1, which obtains OpenAI o3's performance via the self-supervised process reward model (SPRM). Through sharing the backbone network and using task-specific heads for next token prediction and process scoring respectively, SPRM successfully integrates the policy model and process reward model(PRM) into a unified interface without extra process annotation, reducing over 99% PRM parameters for efficient reasoning. Equipped with SPRM, MetaStone-S1 is naturally suitable for test time scaling (TTS), and we provide three reasoning effort modes (low, medium, and high), based on the controllable thinking length. Moreover, we empirically establish a scaling law that reveals the relationship between total thinking computation and TTS performance. Experiments demonstrate that our MetaStone-S1 achieves comparable performance to OpenAI-o3-mini's series with only 32B parameter size. To support the research community, we have open-sourced MetaStone-S1 at https://github.com/MetaStone-AI/MetaStone-S1.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Test-Time Scaling with Reflective Generative Model | TensorX