TensorX
返回文献探索

Paper · arXiv 2312.07000

Alignment for Honesty

Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, Pengfei Liu

13 upvotesDecember 12, 2023arXiv 预印本
AI 摘要

A paper proposes a framework and metrics to enhance the honesty of large language models by training them to accurately assess and reveal their knowledge limits.

alignmenthonestylarge language modelsLLMsfine-tuning

Abstract

Recent research has made significant strides in applying alignment techniques to enhance the helpfulness and harmlessness of large language models (LLMs) in accordance with human intentions. In this paper, we argue for the importance of alignment for honesty, ensuring that LLMs proactively refuse to answer questions when they lack knowledge, while still not being overly conservative. However, a pivotal aspect of alignment for honesty involves discerning the limits of an LLM's knowledge, which is far from straightforward. This challenge demands comprehensive solutions in terms of metric development, benchmark creation, and training methodologies. In this paper, we address these challenges by first establishing a precise problem definition and defining ``honesty'' inspired by the Analects of Confucius. This serves as a cornerstone for developing metrics that effectively measure an LLM's honesty by quantifying its progress post-alignment. Furthermore, we introduce a flexible training framework which is further instantiated by several efficient fine-tuning techniques that emphasize honesty without sacrificing performance on other tasks. Our extensive experiments reveal that these aligned models show a marked increase in honesty, as indicated by our proposed metrics. We open-source a wealth of resources to facilitate future research at https://github.com/GAIR-NLP/alignment-for-honesty, including honesty-aligned models, training and evaluation datasets for honesty alignment, concept glossary, as well as all relevant source code.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号