TensorX
返回文献探索

Paper · arXiv 2602.08321

Improving Data and Reward Design for Scientific Reasoning in Large Language Models

Zijie Chen, Zhenghao Lin, Xiao Liu, Zhenzhong Lan, Yeyun Gong, Peng Cheng

44 upvotesFebruary 9, 2026arXiv 预印本
AI 摘要

A large-scale scientific question dataset and post-training pipeline are developed to improve open-ended science question answering through enhanced data processing and reinforcement learning with rubric-guided evaluation.

large language modelsopen-ended science questionsdata constructionreward designDr. SCI datasetSFTRLexploration-expanding SFTdynamic difficulty curriculumSciRubric-Guided RLscientific reasoningGPQA-diamondGPQA-general

Abstract

Solving open-ended science questions remains challenging for large language models, particularly due to inherently unreliable supervision and evaluation. The bottleneck lies in the data construction and reward design for scientific post-training. We develop a large-scale, systematic data processing pipeline that transforms heterogeneous open-source science data into Dr. SCI dataset, which comprises of 1M questions across eight STEM subjects, with explicit verifiable/open-ended splits, scalable difficulty annotation, and fine-grained rubrics that operationalize evaluation for open-ended answers. Building on this dataset, we propose the Dr. SCI post-training pipeline, which redesigns the standard SFT -> RL workflow through three components: (i) Exploration-Expanding SFT, which broadens the model's reasoning pattern coverage prior to RL; (ii) Dynamic Difficulty Curriculum, which adapts training data to the model's evolving scientific capability; and (iii) SciRubric-Guided RL, which enables stable reinforcement learning on open-ended scientific questions via rubric-based evaluation with explicit answer correctness. Qwen3-4B-Base trained using Dr. SCI pipeline achieves 63.2 on GPQA-diamond and 32.4 on GPQA-general, consistently improves over strong post-trained baselines such as o1-mini and GPT-4o, demonstrating substantial gains in scientific reasoning, especially in open-ended settings.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Improving Data and Reward Design for Scientific Reasoning in Large Language Models | TensorX