Aivora
arXivAI ResearchAdvanced

Rounding in Preconditioner Space: Redesigning 4-bit AdamW Quantization

預調節器空間捨入:重構 4-bit AdamW 優化器狀態量化

2 min read
Rounding in Preconditioner Space: Redesigning 4-bit AdamW Quantization
The 30-second version

Quantizing AdamW optimizer states to 4-bit saves substantial memory, but quantization errors accumulate through moment recurrences and corrupt gradient updates. This study shows that small errors in state space do not guarantee small errors in preconditioner space. To address this, the authors introduce ZIP-SR (zero-inclusive preconditioner-space stochastic rounding) and ZE-EDEN (zero-excluding block rescaling), combined with targeted stochastic rounding for the LM-head during the final 10% of training. Evaluated on 130M to 2.7B parameter models, these methods outperform TorchAO 4-bit AdamW, shrinking the loss gap to FP32 AdamW by up to 70%.

Key points

01

Preconditioner-Space Rounding Analysis

Demonstrates that low quantization error in second-moment state space does not imply low error in preconditioner space, shifting decision boundary to preconditioner coordinates.

02

ZIP-SR & ZE-EDEN Algorithms

ZIP-SR keeps zero and computes stochastic rounding probabilities in preconditioner space, while ZE-EDEN excludes zero and rescales blocks to fix quantization floor bias.

03

Targeted LM-Head First Moment Rounding

Uses NF4 for the first moment, adding targeted stochastic rounding to the LM-head first moment during the final 10% of training.

04

Up to 70% Gap Reduction

Across GPT- and Llama-style pretraining (130M to 2.7B parameters), both methods close the validation loss gap to 32-bit AdamW by up to 70%.

How it works

Preconditioner-Space 4-bit AdamW Quantization Pipeline
update momentaccumulate sqspace transformrounding probcombine updateapply precondGradient g_t1st Moment (NF4 + LateSR)2nd Moment CalculationPreconditioner SpaceMappingZIP-SR / ZE-EDENQuantizationAccurate AdaptiveUpdate

Why it matters

LLM training is often limited by the heavy memory footprint of optimizer states. While 4-bit optimizers (such as TorchAO) reduce memory usage, quantization error propagation worsens final model accuracy. By mathematically reformulating the rounding space, this work recovers performance close to FP32 AdamW without memory trade-offs, enabling memory-efficient yet accurate LLM pretraining and full-parameter SFT.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Memory-efficient pretraining of LLMs ranging from 130M to 2.7B+ parameters
  2. 2Full-parameter Supervised Fine-Tuning (SFT) with significantly reduced VRAM overhead

Limitations & caveats

  • Empirical evaluations are performed up to 2.7B parameters; validation on ultra-large scales (>70B) requires further study.
  • A slight performance gap to full 32-bit AdamW remains inherent to 4-bit state quantization.

Related

Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark
arXivAI Research

Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark

重新審視 METR 時間跨度指標:以統計模型量化 AI 的真實任務能力

Re-analyzing METR benchmark data via splines and item-response theory reveals that AI task difficulty is non-linear, making a jump from 3 to 30 minutes far easier than from 30 minutes to 5 hours.

2 min read