TensorX
返回文献探索

Paper · arXiv 2305.02483

ChatGPT-steered Editing Instructor for Customization of Abstractive Summarization

Wen Xiao, Yujia Xie, Giuseppe Carenini, Pengcheng He

3 upvotesMay 4, 2023arXiv 预印本
AI 摘要

A tri-agent pipeline for enhancing customizability of large language model outputs includes a generator, an instructor, and an editor, with the instructor being optimized through reinforcement learning guided by the editor.

large language modelsChatGPTtri-agent generation pipelinegeneratorinstructoreditorreinforcement learningeditor-steered reinforcement learningabstractive summarization

Abstract

Tailoring outputs of large language models, such as ChatGPT, to specific user needs remains a challenge despite their impressive generation quality. In this paper, we propose a tri-agent generation pipeline consisting of a generator, an instructor, and an editor to enhance the customization of generated outputs. The generator produces an initial output, the user-specific instructor generates editing instructions, and the editor generates a revised output aligned with user preferences. The inference-only large language model (ChatGPT) serves as both the generator and the editor, while a smaller model acts as the user-specific instructor to guide the generation process toward user needs. The instructor is trained using editor-steered reinforcement learning, leveraging feedback from the large-scale editor model to optimize instruction generation. Experimental results on two abstractive summarization datasets demonstrate the effectiveness of our approach in generating outputs that better fulfill user expectations.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ChatGPT-steered Editing Instructor for Customization of Abstractive Summarization | TensorX