TensorX
返回文献探索

Paper · arXiv 2406.14783

Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework

Zackary Rackauckas, Arthur Câmara, Jakub Zavrel

17 upvotesJune 20, 2024arXiv 预印本
AI 摘要

A new evaluation framework using LLMs and an Elo-based system addresses challenges in automating the evaluation of RAG-based QA systems, showing RAG-Fusion outperforms RAG in completeness but not in precision.

Retrieval-Augmented GenerationQ&A systemshallucinationgold standard benchmarksRAG-FusionLarge Language Modelssynthetic queriesLLM-as-a-judgeElo-based competitiondomain expert scoringrelevanceaccuracycompletenessprecisionMRR@5RAGElo

Abstract

Challenges in the automated evaluation of Retrieval-Augmented Generation (RAG) Question-Answering (QA) systems include hallucination problems in domain-specific knowledge and the lack of gold standard benchmarks for company internal tasks. This results in difficulties in evaluating RAG variations, like RAG-Fusion (RAGF), in the context of a product QA task at Infineon Technologies. To solve these problems, we propose a comprehensive evaluation framework, which leverages Large Language Models (LLMs) to generate large datasets of synthetic queries based on real user queries and in-domain documents, uses LLM-as-a-judge to rate retrieved documents and answers, evaluates the quality of answers, and ranks different variants of Retrieval-Augmented Generation (RAG) agents with RAGElo's automated Elo-based competition. LLM-as-a-judge rating of a random sample of synthetic queries shows a moderate, positive correlation with domain expert scoring in relevance, accuracy, completeness, and precision. While RAGF outperformed RAG in Elo score, a significance analysis against expert annotations also shows that RAGF significantly outperforms RAG in completeness, but underperforms in precision. In addition, Infineon's RAGF assistant demonstrated slightly higher performance in document relevance based on MRR@5 scores. We find that RAGElo positively aligns with the preferences of human annotators, though due caution is still required. Finally, RAGF's approach leads to more complete answers based on expert annotations and better answers overall based on RAGElo's evaluation criteria.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework | TensorX