TensorX
返回文献探索

Paper · arXiv 2307.06908

Generating Benchmarks for Factuality Evaluation of Language Models

Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, Yoav Shoham

8 upvotesJuly 13, 2023arXiv 预印本
AI 摘要

FACTOR evaluates language model factuality by transforming a factual corpus into a benchmark of true vs. similar incorrect statements, demonstrating improved accuracy over perplexity in large models and with retrieval augmentation.

FACTORfactual corpusbenchmarklanguage modelfactualityperplexityretrieval augmentationopen-ended generationhuman annotators

Abstract

Before deploying a language model (LM) within a given domain, it is important to measure its tendency to generate factually incorrect information in that domain. Existing factual generation evaluation methods focus on facts sampled from the LM itself, and thus do not control the set of evaluated facts and might under-represent rare and unlikely facts. We propose FACTOR: Factual Assessment via Corpus TransfORmation, a scalable approach for evaluating LM factuality. FACTOR automatically transforms a factual corpus of interest into a benchmark evaluating an LM's propensity to generate true facts from the corpus vs. similar but incorrect statements. We use our framework to create two benchmarks: Wiki-FACTOR and News-FACTOR. We show that: (i) our benchmark scores increase with model size and improve when the LM is augmented with retrieval; (ii) benchmark score correlates with perplexity, but the two metrics do not always agree on model ranking; and (iii) when perplexity and benchmark score disagree, the latter better reflects factuality in open-ended generation, as measured by human annotators. We make our data and code publicly available in https://github.com/AI21Labs/factor.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Generating Benchmarks for Factuality Evaluation of Language Models | TensorX