TensorX
返回文献探索

Paper · arXiv 2411.11767

Drowning in Documents: Consequences of Scaling Reranker Inference

Mathew Jacob, Erik Lindgren, Matei Zaharia, Michael Carbin, Omar Khattab, Andrew Drozdov

19 upvotesNovember 18, 2024arXiv 预印本
AI 摘要

Rerankers, typically cross-encoders, offer diminishing returns and degrade quality when scoring a large number of documents, challenging the assumption that they are consistently more effective.

rerankerscross-encodersinitial IR systemsfull retrievalre-scoringsemantic overlap

Abstract

Rerankers, typically cross-encoders, are often used to re-score the documents retrieved by cheaper initial IR systems. This is because, though expensive, rerankers are assumed to be more effective. We challenge this assumption by measuring reranker performance for full retrieval, not just re-scoring first-stage retrieval. Our experiments reveal a surprising trend: the best existing rerankers provide diminishing returns when scoring progressively more documents and actually degrade quality beyond a certain limit. In fact, in this setting, rerankers can frequently assign high scores to documents with no lexical or semantic overlap with the query. We hope that our findings will spur future research to improve reranking.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号