TensorX
返回文献探索

Paper · arXiv 2501.05707

Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

Vighnesh Subramaniam, Yilun Du, Joshua B. Tenenbaum, Antonio Torralba, Shuang Li, Igor Mordatch

20 upvotesJanuary 10, 2025arXiv 预印本
AI 摘要

A multi-agent society of language models self-improves through interactions, enabling specialization and better performance over rounds of fine-tuning compared to single-agent methods.

large language modelssynthetic dataself-improvementfinetuningmultiagent societyindependent specializationreasoning chains

Abstract

Large language models (LLMs) have achieved remarkable performance in recent years but are fundamentally limited by the underlying training data. To improve models beyond the training data, recent works have explored how LLMs can be used to generate synthetic data for autonomous self-improvement. However, successive steps of self-improvement can reach a point of diminishing returns. In this work, we propose a complementary approach towards self-improvement where finetuning is applied to a multiagent society of language models. A group of language models, all starting from the same base model, are independently specialized by updating each one using data generated through multiagent interactions among the models. By training each model on independent sets of data, we illustrate how this approach enables specialization across models and diversification over the set of models. As a result, our overall system is able to preserve diverse reasoning chains and autonomously improve over many more rounds of fine-tuning than single-agent self-improvement methods. We quantitatively illustrate the efficacy of the approach across a wide suite of reasoning tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains | TensorX