Self-rewarding correction for mathematical reasoning
Wei Xiong, Hanning Zhang, Chenlu Ye +3 authors
Self-rewarding reasoning large language models independently generate and correct their outputs during inference using a two-stage algorithmic framework, enhancing performance without external feedback.