TensorX
返回文献探索

Paper · arXiv 2505.03005

RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale

Daniel Goldstein, Eric Alcaide, Janna Lu, Eugene Cheah

36 upvotesMay 5, 2025arXiv 预印本
AI 摘要

A protocol converts softmax attention transformers into linear attention decoders using minimal tokens while maintaining quality and performance.

softmax attention transformerslinear attention decodersRapid Attention DistillationRADLADSRWKV-variant architecturesQwen2.5 modelsHuggingFaceApache 2.0 licenseQwen License Agreement

Abstract

We present Rapid Attention Distillation to Linear Attention Decoders at Scale (RADLADS), a protocol for rapidly converting softmax attention transformers into linear attention decoder models, along with two new RWKV-variant architectures, and models converted from popular Qwen2.5 open source models in 7B, 32B, and 72B sizes. Our conversion process requires only 350-700M tokens, less than 0.005% of the token count used to train the original teacher models. Converting to our 72B linear attention model costs less than \$2,000 USD at today's prices, yet quality at inference remains close to the original transformer. These models achieve state-of-the-art downstream performance across a set of standard benchmarks for linear attention models of their size. We release all our models on HuggingFace under the Apache 2.0 license, with the exception of our 72B models which are also governed by the Qwen License Agreement. Models at https://huggingface.co/collections/recursal/radlads-6818ee69e99e729ba8a87102 Training Code at https://github.com/recursal/RADLADS-paper

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale | TensorX