TensorX
返回文献探索

Paper · arXiv 2502.17407

Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning

Guijin Son, Jiwoo Hong, Hyunwoo Ko, James Thorne

25 upvotesFebruary 24, 2025arXiv 预印本
AI 摘要

Multilingual math benchmark MCLM tests various test-time scaling methods on multilingual LLMs, showing that performance gains from these methods are limited outside of English.

MCLMOutcome Reward ModelingProcess Reward ModelingBudget ForcingQwen2.5-1.5B MathMR1-1.5Bmultilingual LLMinference FLOPsAIME

Abstract

Scaling pre-training compute has proven effective for achieving mulitlinguality, but does the same hold for test-time scaling? In this work, we introduce MCLM, a multilingual math benchmark featuring competition-level problems in 55 languages. We test three test-time scaling methods-Outcome Reward Modeling (ORM), Process Reward Modeling (ORM), and Budget Forcing (BF)-on both Qwen2.5-1.5B Math and MR1-1.5B, a multilingual LLM we trained for extended reasoning. Our experiments show that using Qwen2.5-1.5B Math with ORM achieves a score of 35.8 on MCLM, while BF on MR1-1.5B attains 35.2. Although "thinking LLMs" have recently garnered significant attention, we find that their performance is comparable to traditional scaling methods like best-of-N once constrained to similar levels of inference FLOPs. Moreover, while BF yields a 20-point improvement on English AIME, it provides only a 1.94-point average gain across other languages-a pattern consistent across the other test-time scaling methods we studied-higlighting that test-time scaling may not generalize as effectively to multilingual tasks. To foster further research, we release MCLM, MR1-1.5B, and evaluation results.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号