TensorX
返回文献探索

Paper · arXiv 2504.13146

Antidistillation Sampling

Yash Savani, Asher Trockman, Zhili Feng, Avi Schwarzschild, Alexander Robey, Marc Finzi, J. Zico Kolter

60 upvotesApril 17, 2025arXiv 预印本
AI 摘要

Antidistillation sampling modifies a model's next-token probability distribution to disrupt the generation of reasoning traces for distillation without affecting model performance.

antidistillation samplingnext-token probability distributionreasoning tracesmodel distillation

Abstract

Frontier models that generate extended reasoning traces inadvertently produce rich token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. Antidistillation sampling provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's practical utility. For further details, see https://antidistillation.com.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Antidistillation Sampling | TensorX