TensorX
返回文献探索

Paper · arXiv 2406.16048

Evaluating D-MERIT of Partial-annotation on Information Retrieval

Royi Rassin, Yaron Fairstein, Oren Kalinsky, Guy Kushilevitz, Nachshon Cohen, Alexander Libov, Yoav Goldberg

36 upvotesJune 23, 2024arXiv 预印本
AI 摘要

Evaluating retrieval models on partially-annotated datasets can lead to misleading rankings, and more comprehensive annotation improves evaluation reliability.

retrieval modelspartially-annotated datasetsrelevant textsfalse negativesD-MERITpassage retrievalWikipediaevaluation settext retrievalresource-efficiencyreliable evaluation

Abstract

Retrieval models are often evaluated on partially-annotated datasets. Each query is mapped to a few relevant texts and the remaining corpus is assumed to be irrelevant. As a result, models that successfully retrieve false negatives are punished in evaluation. Unfortunately, completely annotating all texts for every query is not resource efficient. In this work, we show that using partially-annotated datasets in evaluation can paint a distorted picture. We curate D-MERIT, a passage retrieval evaluation set from Wikipedia, aspiring to contain all relevant passages for each query. Queries describe a group (e.g., ``journals about linguistics'') and relevant passages are evidence that entities belong to the group (e.g., a passage indicating that Language is a journal about linguistics). We show that evaluating on a dataset containing annotations for only a subset of the relevant passages might result in misleading ranking of the retrieval systems and that as more relevant texts are included in the evaluation set, the rankings converge. We propose our dataset as a resource for evaluation and our study as a recommendation for balance between resource-efficiency and reliable evaluation when annotating evaluation sets for text retrieval.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号