TensorX
返回文献探索

Paper · arXiv 2402.14904

Watermarking Makes Language Models Radioactive

Tom Sander, Pierre Fernandez, Alain Durmus, Matthijs Douze, Teddy Furon

23 upvotesFebruary 22, 2024arXiv 预印本
AI 摘要

Watermarked training data can be detected with high confidence in LLMs, making it easier to identify if watermarked outputs were used for fine-tuning compared to conventional methods.

watermarked training datamembership inferencewatermark robustnessfine-tuning

Abstract

This paper investigates the radioactivity of LLM-generated texts, i.e. whether it is possible to detect that such input was used as training data. Conventional methods like membership inference can carry out this detection with some level of accuracy. We show that watermarked training data leaves traces easier to detect and much more reliable than membership inference. We link the contamination level to the watermark robustness, its proportion in the training set, and the fine-tuning process. We notably demonstrate that training on watermarked synthetic instructions can be detected with high confidence (p-value < 1e-5) even when as little as 5% of training text is watermarked. Thus, LLM watermarking, originally designed for detecting machine-generated text, gives the ability to easily identify if the outputs of a watermarked LLM were used to fine-tune another LLM.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Watermarking Makes Language Models Radioactive | TensorX