TensorX
返回文献探索

Paper · arXiv 2502.14382

S*: Test Time Scaling for Code Generation

Dacheng Li, Shiyi Cao, Chengkun Cao, Xiuyu Li, Shangyin Tan, Kurt Keutzer, Jiarong Xing, Joseph E. Gonzalez, Ion Stoica

63 upvotesFebruary 20, 2025arXiv 预印本
AI 摘要

A hybrid test-time scaling framework improves code generation coverage and accuracy across various models and domains.

hybrid test-time scaling frameworkparallel scalingsequential scalingselection mechanismdistinguishing inputsexecution-grounded informationLarge Language ModelsLarge Reasoning ModelsLiveCodeBenchDeepSeek

Abstract

Increasing test-time compute for LLMs shows promise across domains but remains underexplored in code generation, despite extensive study in math. In this paper, we propose S*, the first hybrid test-time scaling framework that substantially improves the coverage and selection accuracy of generated code. S* extends the existing parallel scaling paradigm with sequential scaling to push performance boundaries. It further leverages a novel selection mechanism that adaptively generates distinguishing inputs for pairwise comparison, combined with execution-grounded information to robustly identify correct solutions. We evaluate across 12 Large Language Models and Large Reasoning Model and show: (1) S* consistently improves performance across model families and sizes, enabling a 3B model to outperform GPT-4o-mini; (2) S* enables non-reasoning models to surpass reasoning models - GPT-4o-mini with S* outperforms o1-preview by 3.7% on LiveCodeBench; (3) S* further boosts state-of-the-art reasoning models - DeepSeek-R1-Distill-Qwen-32B with S* achieves 85.7% on LiveCodeBench, approaching o1 (high) at 88.5%. Code will be available under https://github.com/NovaSky-AI/SkyThought.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号