TensorX
返回文献探索

Paper · arXiv 2409.05806

Benchmarking Chinese Knowledge Rectification in Large Language Models

Tianhe Lu, Jizhan Fang, Yunzhi Yao, Xin Xu, Ningyu Zhang, Huajun Chen

14 upvotesSeptember 9, 2024arXiv 预印本
AI 摘要

A benchmark dataset (CKnowEdit) is introduced to address knowledge gaps in LLMs when generating Chinese content, highlighting areas requiring improvement in knowledge editing techniques.

Large Language Models (LLMs)hallucinationsChinese knowledgeCKnowEditknowledge editingclassical textsidiomsBaidu Tieba Ruozhibapolyphonyantithesislogical constructs

Abstract

While Large Language Models (LLMs) exhibit remarkable generative capabilities, they are not without flaws, particularly in the form of hallucinations. This issue is even more pronounced when LLMs are applied to specific languages and domains. For example, LLMs may generate nonsense information when handling Chinese ancient poetry, proverbs, or idioms, owing to the lack of specific knowledge. To this end, this paper introduces a benchmark for rectifying Chinese knowledge in LLMs via knowledge editing. Specifically, we introduce a new Chinese dataset, CKnowEdit, by collecting seven type of knowledge from various sources, including classical texts, idioms, and content from Baidu Tieba Ruozhiba, thereby accounting for the unique polyphony, antithesis, and logical constructs inherent in the Chinese language. Through the analysis of this dataset, we uncover the challenges faced by current LLMs in mastering Chinese. Furthermore, our evaluation of state-of-the-art knowledge editing techniques on this dataset unveil the substantial scope for advancement in the rectification of Chinese knowledge. Code and dataset are available at https://github.com/zjunlp/EasyEdit.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号