TensorX
返回文献探索

Paper · arXiv 2412.11919

RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation

Xiaoxi Li, Jiajie Jin, Yujia Zhou, Yongkang Wu, Zhonghua Li, Qi Ye, Zhicheng Dou

36 upvotesDecember 16, 2024arXiv 预印本
AI 摘要

RetroLLM is a unified framework that integrates retrieval and generation for LLMs, using constraints and decoding strategies to enhance performance and accuracy in evidence generation.

LLMshallucinationsretrieval-augmented generation (RAG)unified frameworkconstrained decodinghierarchical FM-Index constraintsforward-looking constrained decoding strategyextensive experimentsopen-domain QA datasets

Abstract

Large language models (LLMs) exhibit remarkable generative capabilities but often suffer from hallucinations. Retrieval-augmented generation (RAG) offers an effective solution by incorporating external knowledge, but existing methods still face several limitations: additional deployment costs of separate retrievers, redundant input tokens from retrieved text chunks, and the lack of joint optimization of retrieval and generation. To address these issues, we propose RetroLLM, a unified framework that integrates retrieval and generation into a single, cohesive process, enabling LLMs to directly generate fine-grained evidence from the corpus with constrained decoding. Moreover, to mitigate false pruning in the process of constrained evidence generation, we introduce (1) hierarchical FM-Index constraints, which generate corpus-constrained clues to identify a subset of relevant documents before evidence generation, reducing irrelevant decoding space; and (2) a forward-looking constrained decoding strategy, which considers the relevance of future sequences to improve evidence accuracy. Extensive experiments on five open-domain QA datasets demonstrate RetroLLM's superior performance across both in-domain and out-of-domain tasks. The code is available at https://github.com/sunnynexus/RetroLLM.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号