STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸
Linear attention models replace growing KV caches with fixed-size recurrent states, which still pose memory bottlenecks during high-concurrency serving. Directly quantizing these states causes severe error accumulation over time. STEPQuant solves this by analyzing errors across two dimensions: temporally (allocating precision based on memory lifetime) and spatially (fitting scales based on key-row and value-column importance). It closely matches FP32 accuracy at 6-bit, achieves over 5x state compression, and reduces total serving memory by up to 68.7% when integrated into SGLang.
Key points
Temporal Error Control
Mitigates error propagation in long-lived memory over time steps by dynamically allocating precision based on lifetime.
Spatial Sensitivity Fit
Jointly fits key-row and value-column scales to adapt to non-uniform state magnitudes and differing output impacts.
Massive Memory Savings
Integrated into SGLang with optimized GPU kernels, delivering over 5x compression and up to 68.7% serving memory reduction.
High Low-Bit Accuracy
On Qwen and Kimi linear models, the 6-bit config matches FP32 accuracy, and the 4-bit config outperforms uniform INT8.
How it works
Why it matters
As LLMs scale to longer contexts and higher concurrency, the memory footprint of states becomes critical. While linear attention offers fixed-size states, quantizing them has been hindered by recurrent error accumulation. STEPQuant unlocks effective state quantization without accuracy degradation, drastically reducing serving costs and paving the way for highly efficient, long-context linear attention LLM deployments.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1High-concurrency serving optimization for linear attention LLMs
- 2Memory footprint reduction in long-context generation tasks
Limitations & caveats
- Primarily designed for Delta-rule recurrent states; generalization to other non-linear attention mechanisms remains to be verified.
- Requires dedicated operator support in inference engines like SGLang to fully unlock its serving performance benefits.
Related
LeapQuant: Near-Lossless 8-Bit Recurrent State Quantization for Linear Attention LLMs
LeapQuant:突破線性注意力瓶頸,實現近乎無損的 8-bit 遞迴狀態量化
LeapQuant is a training-free quantization method that achieves near-lossless 8-bit recurrent state quantization for linear attention, delivering up to 1.47x end-to-end inference speedup while preserving FP32 accuracy.
Telescopic Language Models: One Training Run for Endless Compute Budgets
伸縮自如的語言模型:單次訓練即可適應多種運算資源預算的 Telescopic LM
This research introduces Telescopic Language Models (TLM), which use stochastic prefix supervision to enable a single Transformer to act as a valid language model at any layer depth, serving diverse compute budgets from a single training run.
Minimally Invasive Steering of LMs: Optimizing Rewards Without Quality Degradation
微創型語言模型導向技術:利用 MISVO 在不損害生成品質下優化輸出
This paper introduces MISVO, a minimally invasive steering method that uses local KL geometry to optimize LLM outputs for test-time rewards without parameter updates or quality degradation.