LeapQuant: Near-Lossless 8-Bit Recurrent State Quantization for Linear Attention LLMs
LeapQuant:突破線性注意力瓶頸,實現近乎無損的 8-bit 遞迴狀態量化
While hybrid LLMs utilize linear attention to compress context into a fixed recurrent state, the frequent read/write cycles of this state bottleneck inference. LeapQuant introduces a training-free 8-bit quantization framework to address this. It implements "per-window quantization" to jump over token windows and reduce error accumulation, keeps prominent outliers as high-precision "Compensator Tokens", and smooths the remaining residuals. Evaluated on Qwen, Kimi, and GLM models, LeapQuant delivers up to 1.47x end-to-end speedups on modern GPUs (including NVIDIA B200 and RTX 5090) with near-FP32 accuracy.
Key points
Per-Window Quantization
Leaps over a window of tokens and quantizes the state only once at its end, computing in-window outputs using fixed low-bit states and high-precision buffered updates.
Compensator Tokens
Retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens to minimize quantization errors.
Residual Smoothing
Smooths the remaining residual matrix before quantization to further compress the quantization error of non-outlier values.
Significant Speedups
Achieves average speedups of 2.05x to 3.70x at the kernel level and 1.47x for end-to-end inference on high-end GPUs.
How it works
Why it matters
As long-context demands surge, linear attention is vital for efficiency, but quantizing its recurrent state has remained a bottleneck. LeapQuant proves that recurrent states can be safely quantized to 8-bit without retraining. This paves the way for deploying next-gen hybrid LLMs efficiently on both consumer GPUs (e.g., RTX 5090) and enterprise hardware (e.g., B200), significantly slashing memory and compute costs.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Hardware deployment and acceleration for long-context hybrid LLMs such as Kimi Delta Attention and Gated DeltaNet
- 2Running large language models on memory-constrained edge devices or consumer GPUs like the RTX 5090
Limitations & caveats
- Specifically designed for linear attention and hybrid models, and may not directly apply to traditional Softmax attention architectures
- Still requires keeping buffered updates in high precision within the local window, which may consume memory under certain hyperparameters
Related
STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸
STEPQuant is a spatial-temporal post-training quantization framework for Delta-rule recurrent states, achieving over 5x state compression and reducing serving memory by up to 68.7% with minimal accuracy loss.
Telescopic Language Models: One Training Run for Endless Compute Budgets
伸縮自如的語言模型:單次訓練即可適應多種運算資源預算的 Telescopic LM
This research introduces Telescopic Language Models (TLM), which use stochastic prefix supervision to enable a single Transformer to act as a valid language model at any layer depth, serving diverse compute budgets from a single training run.
Minimally Invasive Steering of LMs: Optimizing Rewards Without Quality Degradation
微創型語言模型導向技術:利用 MISVO 在不損害生成品質下優化輸出
This paper introduces MISVO, a minimally invasive steering method that uses local KL geometry to optimize LLM outputs for test-time rewards without parameter updates or quality degradation.