TensorX
返回文献探索

Paper · arXiv 2411.08147

Large Language Models Can Self-Improve in Long-context Reasoning

Siheng Li, Cheng Yang, Zesen Cheng, Lemao Liu, Mo Yu, Yujiu Yang, Wai Lam

65 upvotesNovember 12, 2024arXiv 预印本
AI 摘要

A novel method enables LLMs to self-improve in long-context reasoning through sampling, scoring with Minimum Bayes Risk, and supervised fine-tuning, outperforming expert-annotated data approaches.

LLMslong-context reasoningfine-tuningsynthetic dataGPT-4Minimum Bayes Risksupervised fine-tuningpreference optimizationLlama-3.1-8B-Instruct

Abstract

Large language models (LLMs) have achieved substantial progress in processing long contexts but still struggle with long-context reasoning. Existing approaches typically involve fine-tuning LLMs with synthetic data, which depends on annotations from human experts or advanced models like GPT-4, thus restricting further advancements. To address this issue, we investigate the potential for LLMs to self-improve in long-context reasoning and propose \ours, an approach specifically designed for this purpose. This approach is straightforward: we sample multiple outputs for each question, score them with Minimum Bayes Risk, and then apply supervised fine-tuning or preference optimization based on these outputs. Extensive experiments on several leading LLMs demonstrate the effectiveness of \ours, with an absolute improvement of 4.2 points for Llama-3.1-8B-Instruct. Furthermore, \ours achieves superior performance compared to prior approaches that depend on data produced by human experts or advanced models. We anticipate that this work will open new avenues for self-improvement techniques in long-context scenarios, which are essential for the continual advancement of LLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号