AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
Zihan Liu, Zhuolin Yang, Yang Chen +4 authors
Combining supervised fine-tuning and reinforcement learning enhances reasoning models, especially when optimizing sampling temperature and leveraging strong initial fine-tuning, as demonstrated by the improved AceReason-Nemotron-1.1 model.