TensorX
返回文献探索

Paper · arXiv 2308.04623

Accelerating LLM Inference with Staged Speculative Decoding

Benjamin Spector, Chris Re

26 upvotesAugust 8, 2023arXiv 预印本
AI 摘要

A staged speculative decoding algorithm enhances LLM inference efficiency in small-batch, on-device settings by restructuring speculative batches and adding a second decoding stage.

large language modelsstaged speculative decodingspeculative decodingtree structureGPT-2-Lparameter-efficient fine-tuning

Abstract

Recent advances with large language models (LLM) illustrate their diverse capabilities. We propose a novel algorithm, staged speculative decoding, to accelerate LLM inference in small-batch, on-device scenarios. We address the low arithmetic intensity of small-batch inference by improving upon previous work in speculative decoding. First, we restructure the speculative batch as a tree, which reduces generation costs and increases the expected tokens per batch. Second, we add a second stage of speculative decoding. Taken together, we reduce single-batch decoding latency by 3.16x with a 762M parameter GPT-2-L model while perfectly preserving output quality.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号