TensorX
返回文献探索

Paper · arXiv 2607.27951

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu

7 upvotesJuly 30, 2026arXiv 预印本
AI 摘要

Separating model capabilities from downstream-use evidence reveals a safety trilemma among utility, reliability, and openness, and shows that trusted credentials can improve safeguards by adding hard-to-copy predictive signals.

large language model safeguardsdual-use taskstrusted credentialssafety trilemmaopen accessadaptive attacks

Abstract

Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs | TensorX