TensorX
返回文献探索

Paper · arXiv 2312.02969

Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models

Xinyu Zhang, Sebastian Hofstätter, Patrick Lewis, Raphael Tang, Jimmy Lin

14 upvotesDecember 5, 2023arXiv 预印本
AI 摘要

The study develops effective listwise rerankers independent of GPT models, surpassing GPT-3.5 and achieving near-GPT-4 performance, and highlights the need for high-quality listwise ranking datasets.

listwise rerankerslarge language modelsLLMGPT modelspassage retrievallist se rerankerpointwise rankinghigh-quality listwise ranking datahuman-annotated listwise data resources

Abstract

Listwise rerankers based on large language models (LLM) are the zero-shot state-of-the-art. However, current works in this direction all depend on the GPT models, making it a single point of failure in scientific reproducibility. Moreover, it raises the concern that the current research findings only hold for GPT models but not LLM in general. In this work, we lift this pre-condition and build for the first time effective listwise rerankers without any form of dependency on GPT. Our passage retrieval experiments show that our best list se reranker surpasses the listwise rerankers based on GPT-3.5 by 13% and achieves 97% effectiveness of the ones built on GPT-4. Our results also show that the existing training datasets, which were expressly constructed for pointwise ranking, are insufficient for building such listwise rerankers. Instead, high-quality listwise ranking data is required and crucial, calling for further work on building human-annotated listwise data resources.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号