TensorX
返回文献探索

Paper · arXiv 2502.09604

SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models

Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Zejiang Shen, Zhaofeng Wu, Hu Xu, Xi Victoria Lin, James Glass, Shang-Wen Li, Wen-tau Yih

36 upvotesFebruary 13, 2025arXiv 预印本
AI 摘要

SelfCite is a self-supervised method that uses context ablation to align LLMs for generating high-quality, sentence-level citations, improving citation quality and F1 scores on long-form question answering tasks.

self-supervisedLLMssentence-level citationscontext ablationreward signalbest-of-N samplingpreference optimizationcitation F1LongBench-Cite

Abstract

We introduce SelfCite, a novel self-supervised approach that aligns LLMs to generate high-quality, fine-grained, sentence-level citations for the statements in their generated responses. Instead of only relying on costly and labor-intensive annotations, SelfCite leverages a reward signal provided by the LLM itself through context ablation: If a citation is necessary, removing the cited text from the context should prevent the same response; if sufficient, retaining the cited text alone should preserve the same response. This reward can guide the inference-time best-of-N sampling strategy to improve citation quality significantly, as well as be used in preference optimization to directly fine-tune the models for generating better citations. The effectiveness of SelfCite is demonstrated by increasing citation F1 up to 5.3 points on the LongBench-Cite benchmark across five long-form question answering tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models | TensorX