TensorX
返回文献探索

Paper · arXiv 2608.07594

Scaling Inherently Interpretable Language Models

Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail, Giang Nguyen, Isaac Plant, Muawiz Chaudhary, Nathaniel Monson, Saqib Azim, Zhichen Guo, Julius Adebayo

21 upvotesAugust 6, 2026arXiv 预印本
AI 摘要

Integrating interpretability as a training constraint yields scalable, disentangled representations that enable attribution, retrieval, and steering without retraining.

autoregressive language modelsdiffusion language modelscausal attention maskdisentangled representationsconcept attributionfeature attributiontraining data retrievalconcept steeringscaling paradigm

Abstract

Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on both autoregressive and diffusion language models, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts with scale. We instantiate the training-time recipe with Steerling-8B, a diffusion language model with a causal attention mask. For any group of generated tokens, Steerling-8B attributes the output to relevant input tokens, human-understandable concepts, and training data. This enables closed-loop intervention: diagnose an output through its concept or feature attribution, retrieve similar training data, and correct the behavior through concept steering without retraining. Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号