TensorX
返回文献探索

Paper · arXiv 2503.03601

Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders

Kristian Kuznetsov, Laida Kushnareva, Polina Druzhinina, Anton Razzhigaev, Anastasia Voznyuk, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov

234 upvotesMarch 5, 2025arXiv 预印本
AI 摘要

Sparse Autoencoders enhance interpretability in Artificial Text Detection by extracting distinctive features from LLM outputs, providing insights into differences from human-written texts.

Sparse Autoencodersresidual streaminterpretable featuresdomain-specific statisticssteering approachmodel-specific statistics

Abstract

Artificial Text Detection (ATD) is becoming increasingly important with the rise of advanced Large Language Models (LLMs). Despite numerous efforts, no single algorithm performs consistently well across different types of unseen text or guarantees effective generalization to new LLMs. Interpretability plays a crucial role in achieving this goal. In this study, we enhance ATD interpretability by using Sparse Autoencoders (SAE) to extract features from Gemma-2-2b residual stream. We identify both interpretable and efficient features, analyzing their semantics and relevance through domain- and model-specific statistics, a steering approach, and manual or LLM-based interpretation. Our methods offer valuable insights into how texts from various models differ from human-written content. We show that modern LLMs have a distinct writing style, especially in information-dense domains, even though they can produce human-like outputs with personalized prompts.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders | TensorX