Aivora
arXivLLMAdvanced

LeapQuant: Near-Lossless 8-Bit Recurrent State Quantization for Linear Attention LLMs

LeapQuant:突破線性注意力瓶頸,實現近乎無損的 8-bit 遞迴狀態量化

2 min read
LeapQuant: Near-Lossless 8-Bit Recurrent State Quantization for Linear Attention LLMs
The 30-second version

While hybrid LLMs utilize linear attention to compress context into a fixed recurrent state, the frequent read/write cycles of this state bottleneck inference. LeapQuant introduces a training-free 8-bit quantization framework to address this. It implements "per-window quantization" to jump over token windows and reduce error accumulation, keeps prominent outliers as high-precision "Compensator Tokens", and smooths the remaining residuals. Evaluated on Qwen, Kimi, and GLM models, LeapQuant delivers up to 1.47x end-to-end speedups on modern GPUs (including NVIDIA B200 and RTX 5090) with near-FP32 accuracy.

Key points

01

Per-Window Quantization

Leaps over a window of tokens and quantizes the state only once at its end, computing in-window outputs using fixed low-bit states and high-precision buffered updates.

02

Compensator Tokens

Retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens to minimize quantization errors.

03

Residual Smoothing

Smooths the remaining residual matrix before quantization to further compress the quantization error of non-outlier values.

04

Significant Speedups

Achieves average speedups of 2.05x to 3.70x at the kernel level and 1.47x for end-to-end inference on high-end GPUs.

How it works

LeapQuant Recurrent State Quantization Flow
Per-window processIsolate extreme valuesProcess remaining matrixMergeMergeOutput new stateInput Token WindowLow-bit State + FPBuffered UpdateExtract Outliers asCompensatorsResidual Smoothing8-bit QuantizationUpdated Recurrent State

Why it matters

As long-context demands surge, linear attention is vital for efficiency, but quantizing its recurrent state has remained a bottleneck. LeapQuant proves that recurrent states can be safely quantized to 8-bit without retraining. This paves the way for deploying next-gen hybrid LLMs efficiently on both consumer GPUs (e.g., RTX 5090) and enterprise hardware (e.g., B200), significantly slashing memory and compute costs.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Hardware deployment and acceleration for long-context hybrid LLMs such as Kimi Delta Attention and Gated DeltaNet
  2. 2Running large language models on memory-constrained edge devices or consumer GPUs like the RTX 5090

Limitations & caveats

  • Specifically designed for linear attention and hybrid models, and may not directly apply to traditional Softmax attention architectures
  • Still requires keeping buffered updates in high precision within the local window, which may consume memory under certain hyperparameters

Related

STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
arXivLLM

STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States

STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸

STEPQuant is a spatial-temporal post-training quantization framework for Delta-rule recurrent states, achieving over 5x state compression and reducing serving memory by up to 68.7% with minimal accuracy loss.

2 min read
Telescopic Language Models: One Training Run for Endless Compute Budgets
arXivLLM

Telescopic Language Models: One Training Run for Endless Compute Budgets

伸縮自如的語言模型:單次訓練即可適應多種運算資源預算的 Telescopic LM

This research introduces Telescopic Language Models (TLM), which use stochastic prefix supervision to enable a single Transformer to act as a valid language model at any layer depth, serving diverse compute budgets from a single training run.

2 min read