TensorX
返回文献探索

Paper · arXiv 2505.21115

Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

Sergey Pletenev, Maria Marina, Nikolay Ivanov, Daria Galimzianova, Nikita Krayko, Mikhail Salnikov, Vasily Konovalov, Alexander Panchenko, Viktor Moskvoretskii

144 upvotesMay 27, 2025arXiv 预印本
AI 摘要

EverGreenQA, a multilingual QA dataset with evergreen labels, is introduced to benchmark LLMs on temporality encoding and assess their performance through verbalized judgments and uncertainty signals.

Large Language ModelsQAevergreenmutabletemporalityMultilingual QA datasetEG-E5lightweight multilingual classifierSoTA performanceself-knowledge estimationfiltering QA datasetsGPT-4o retrieval behavior

Abstract

Large Language Models (LLMs) often hallucinate in question answering (QA) tasks. A key yet underexplored factor contributing to this is the temporality of questions -- whether they are evergreen (answers remain stable over time) or mutable (answers change). In this work, we introduce EverGreenQA, the first multilingual QA dataset with evergreen labels, supporting both evaluation and training. Using EverGreenQA, we benchmark 12 modern LLMs to assess whether they encode question temporality explicitly (via verbalized judgments) or implicitly (via uncertainty signals). We also train EG-E5, a lightweight multilingual classifier that achieves SoTA performance on this task. Finally, we demonstrate the practical utility of evergreen classification across three applications: improving self-knowledge estimation, filtering QA datasets, and explaining GPT-4o retrieval behavior.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA | TensorX