TensorX
返回文献探索

Paper · arXiv 2408.15666

StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements

Jillian Fisher, Skyler Hallinan, Ximing Lu, Mitchell Gordon, Zaid Harchaoui, Yejin Choi

12 upvotesAugust 28, 2024arXiv 预印本
AI 摘要

StyleRemix is an adaptive and interpretable authorship obfuscation method using pre-trained LoRA modules to perturb specific stylistic elements, outperforming state-of-the-art methods and larger LLMs.

StyleRemixLow Rank Adaptation (LoRA)stylistic elementsauthorship obfuscation

Abstract

Authorship obfuscation, rewriting a text to intentionally obscure the identity of the author, is an important but challenging task. Current methods using large language models (LLMs) lack interpretability and controllability, often ignoring author-specific stylistic features, resulting in less robust performance overall. To address this, we develop StyleRemix, an adaptive and interpretable obfuscation method that perturbs specific, fine-grained style elements of the original input text. StyleRemix uses pre-trained Low Rank Adaptation (LoRA) modules to rewrite an input specifically along various stylistic axes (e.g., formality and length) while maintaining low computational cost. StyleRemix outperforms state-of-the-art baselines and much larger LLMs in a variety of domains as assessed by both automatic and human evaluation. Additionally, we release AuthorMix, a large set of 30K high-quality, long-form texts from a diverse set of 14 authors and 4 domains, and DiSC, a parallel corpus of 1,500 texts spanning seven style axes in 16 unique directions

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号