TensorX
返回文献探索

Paper · arXiv 2407.10969

Q-Sparse: All Large Language Models can be Fully Sparsely-Activated

Hongyu Wang, Shuming Ma, Ruiping Wang, Furu Wei

23 upvotesJuly 15, 2024arXiv 预印本
AI 摘要

Q-Sparse enhances sparsely-activated large language models for efficient inference across various training and deployment scenarios, using top-K sparsification and straight-through-estimators.

sparsely-activated large language modelsQ-Sparsetop-K sparsificationstraight-through-estimatorinference-optimal scaling lawtraining-from-scratchcontinue-trainingfinetuningBitNet b1.58MoE

Abstract

We introduce, Q-Sparse, a simple yet effective approach to training sparsely-activated large language models (LLMs). Q-Sparse enables full sparsity of activations in LLMs which can bring significant efficiency gains in inference. This is achieved by applying top-K sparsification to the activations and the straight-through-estimator to the training. The key results from this work are, (1) Q-Sparse can achieve results comparable to those of baseline LLMs while being much more efficient at inference time; (2) We present an inference-optimal scaling law for sparsely-activated LLMs; (3) Q-Sparse is effective in different settings, including training-from-scratch, continue-training of off-the-shelf LLMs, and finetuning; (4) Q-Sparse works for both full-precision and 1-bit LLMs (e.g., BitNet b1.58). Particularly, the synergy of BitNet b1.58 and Q-Sparse (can be equipped with MoE) provides the cornerstone and a clear path to revolutionize the efficiency, including cost and energy consumption, of future LLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated | TensorX