TensorX
返回文献探索

Paper · arXiv 2409.05152

OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs

Jintian Zhang, Cheng Peng, Mengshu Sun, Xiang Chen, Lei Liang, Zhiqiang Zhang, Jun Zhou, Huajun Chen, Ningyu Zhang

31 upvotesSeptember 8, 2024arXiv 预印本
AI 摘要

OneGen is a framework that integrates retrieval and generation in a single pass within Large Language Models, improving retrieval performance without compromising generative capabilities.

Large Language ModelsOne-pass Generationretrieval tokensautoregressive generationunified forward passRAGEntity Linkingvector retrieval

Abstract

Despite the recent advancements in Large Language Models (LLMs), which have significantly enhanced the generative capabilities for various NLP tasks, LLMs still face limitations in directly handling retrieval tasks. However, many practical applications demand the seamless integration of both retrieval and generation. This paper introduces a novel and efficient One-pass Generation and retrieval framework (OneGen), designed to improve LLMs' performance on tasks that require both generation and retrieval. The proposed framework bridges the traditionally separate training approaches for generation and retrieval by incorporating retrieval tokens generated autoregressively. This enables a single LLM to handle both tasks simultaneously in a unified forward pass. We conduct experiments on two distinct types of composite tasks, RAG and Entity Linking, to validate the pluggability, effectiveness, and efficiency of OneGen in training and inference. Furthermore, our results show that integrating generation and retrieval within the same context preserves the generative capabilities of LLMs while improving retrieval performance. To the best of our knowledge, OneGen is the first to enable LLMs to conduct vector retrieval during the generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs | TensorX