Aivora
arXivLLMAdvanced

STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States

STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸

2 min read
STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
The 30-second version

Linear attention models replace growing KV caches with fixed-size recurrent states, which still pose memory bottlenecks during high-concurrency serving. Directly quantizing these states causes severe error accumulation over time. STEPQuant solves this by analyzing errors across two dimensions: temporally (allocating precision based on memory lifetime) and spatially (fitting scales based on key-row and value-column importance). It closely matches FP32 accuracy at 6-bit, achieves over 5x state compression, and reduces total serving memory by up to 68.7% when integrated into SGLang.

Key points

01

Temporal Error Control

Mitigates error propagation in long-lived memory over time steps by dynamically allocating precision based on lifetime.

02

Spatial Sensitivity Fit

Jointly fits key-row and value-column scales to adapt to non-uniform state magnitudes and differing output impacts.

03

Massive Memory Savings

Integrated into SGLang with optimized GPU kernels, delivering over 5x compression and up to 68.7% serving memory reduction.

04

High Low-Bit Accuracy

On Qwen and Kimi linear models, the 6-bit config matches FP32 accuracy, and the 4-bit config outperforms uniform INT8.

How it works

STEPQuant Spatial-Temporal Quantization Workflow
Lifetime checkMagnitude checkBit allocationApply scalesKernel accelRecurrent StateTemporal AnalysisSpatial AnalysisPrecision AllocScale FittingQuantized StateSGLang Engine

Why it matters

As LLMs scale to longer contexts and higher concurrency, the memory footprint of states becomes critical. While linear attention offers fixed-size states, quantizing them has been hindered by recurrent error accumulation. STEPQuant unlocks effective state quantization without accuracy degradation, drastically reducing serving costs and paving the way for highly efficient, long-context linear attention LLM deployments.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1High-concurrency serving optimization for linear attention LLMs
  2. 2Memory footprint reduction in long-context generation tasks

Limitations & caveats

  • Primarily designed for Delta-rule recurrent states; generalization to other non-linear attention mechanisms remains to be verified.
  • Requires dedicated operator support in inference engines like SGLang to fully unlock its serving performance benefits.

Related

Telescopic Language Models: One Training Run for Endless Compute Budgets
arXivLLM

Telescopic Language Models: One Training Run for Endless Compute Budgets

伸縮自如的語言模型:單次訓練即可適應多種運算資源預算的 Telescopic LM

This research introduces Telescopic Language Models (TLM), which use stochastic prefix supervision to enable a single Transformer to act as a valid language model at any layer depth, serving diverse compute budgets from a single training run.

2 min read