TensorX
返回文献探索

Paper · arXiv 2409.00729

ContextCite: Attributing Model Generation to Context

Benjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander Madry

14 upvotesSeptember 1, 2024arXiv 预印本
AI 摘要

ContextCite is a method for attributing parts of the context used by language models to generate statements, aiding in verification, improving response quality, and detecting poisoning attacks.

context attributionlanguage modelsContextCitegenerated statementsresponse qualitypoisoning attacks

Abstract

How do language models use information provided as context when generating a response? Can we infer whether a particular generated statement is actually grounded in the context, a misinterpretation, or fabricated? To help answer these questions, we introduce the problem of context attribution: pinpointing the parts of the context (if any) that led a model to generate a particular statement. We then present ContextCite, a simple and scalable method for context attribution that can be applied on top of any existing language model. Finally, we showcase the utility of ContextCite through three applications: (1) helping verify generated statements (2) improving response quality by pruning the context and (3) detecting poisoning attacks. We provide code for ContextCite at https://github.com/MadryLab/context-cite.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ContextCite: Attributing Model Generation to Context | TensorX