Rounding in Preconditioner Space: Redesigning 4-bit AdamW Quantization
預調節器空間捨入:重構 4-bit AdamW 優化器狀態量化
Quantizing AdamW optimizer states to 4-bit saves substantial memory, but quantization errors accumulate through moment recurrences and corrupt gradient updates. This study shows that small errors in state space do not guarantee small errors in preconditioner space. To address this, the authors introduce ZIP-SR (zero-inclusive preconditioner-space stochastic rounding) and ZE-EDEN (zero-excluding block rescaling), combined with targeted stochastic rounding for the LM-head during the final 10% of training. Evaluated on 130M to 2.7B parameter models, these methods outperform TorchAO 4-bit AdamW, shrinking the loss gap to FP32 AdamW by up to 70%.
Key points
Preconditioner-Space Rounding Analysis
Demonstrates that low quantization error in second-moment state space does not imply low error in preconditioner space, shifting decision boundary to preconditioner coordinates.
ZIP-SR & ZE-EDEN Algorithms
ZIP-SR keeps zero and computes stochastic rounding probabilities in preconditioner space, while ZE-EDEN excludes zero and rescales blocks to fix quantization floor bias.
Targeted LM-Head First Moment Rounding
Uses NF4 for the first moment, adding targeted stochastic rounding to the LM-head first moment during the final 10% of training.
Up to 70% Gap Reduction
Across GPT- and Llama-style pretraining (130M to 2.7B parameters), both methods close the validation loss gap to 32-bit AdamW by up to 70%.
How it works
Why it matters
LLM training is often limited by the heavy memory footprint of optimizer states. While 4-bit optimizers (such as TorchAO) reduce memory usage, quantization error propagation worsens final model accuracy. By mathematically reformulating the rounding space, this work recovers performance close to FP32 AdamW without memory trade-offs, enabling memory-efficient yet accurate LLM pretraining and full-parameter SFT.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Memory-efficient pretraining of LLMs ranging from 130M to 2.7B+ parameters
- 2Full-parameter Supervised Fine-Tuning (SFT) with significantly reduced VRAM overhead
Limitations & caveats
- Empirical evaluations are performed up to 2.7B parameters; validation on ultra-large scales (>70B) requires further study.
- A slight performance gap to full 32-bit AdamW remains inherent to 4-bit state quantization.
Related
Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark
重新審視 METR 時間跨度指標:以統計模型量化 AI 的真實任務能力
Re-analyzing METR benchmark data via splines and item-response theory reveals that AI task difficulty is non-linear, making a jump from 3 to 30 minutes far easier than from 30 minutes to 5 hours.
Bi-FORK: Generative Modeling for High-Dimensional Bifurcating Physical Systems
Bi-FORK:高維分歧物理系統的生成式建模框架
Bi-FORK combines latent flow matching with repulsion-guided sampling to capture one-to-many physical bifurcations, scaling to high-dimensional systems with up to 260,000 points.
One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
單一區塊實現多重深度:具備深度編程專家庫的循環 Vision Transformer
The paper introduces reViT, which reuses a single Transformer block recurrently by dynamically merging shared expert weights per depth, matching full-depth ViT accuracy with roughly 70% fewer stored parameters.