TACO Optimizer: Unleashing Full-Parameter 32B LLM Fine-Tuning on a Single GPU
TACO 最佳化器:將 32B 大模型全參數微調帶入單張 GPU 的極簡幾何學
Full-parameter LLM fine-tuning is heavily bottlenecked by the massive state memory of optimizers like AdamW. This paper introduces TACO, which computes updates by selecting only the sign of the absolute-maximum gradient entry in each column of the weight matrix. On OPT-13B, TACO reduces persistent optimizer states by 174x (from 27.7 GB to 0.16 GB) and peak training memory by 2.9x compared to 8-bit AdamW. Maintaining comparable accuracy and runtime, TACO enables full-parameter fine-tuning of 30-32B models on a single 80GB H100 GPU.
Key points
Negligible Optimizer State
Reduces persistent optimizer state from 27.7 GB to 0.16 GB for OPT-13B, representing a 174x memory saving compared to 8-bit AdamW.
Column-wise One-sparse Geometry
Computes exact steepest descent by extracting only the sign of the largest magnitude entry in each column of 2D weight matrices.
Single-GPU Scaling Boost
Enables full-parameter fine-tuning of 30-32B parameter models across multiple model families on a single 80 GB H100 GPU.
Comparable Fine-tuning Accuracy
Overcomes geometric discrepancy issues of prior approaches like Muon, ensuring convergence quality and accuracy on par with AdamW.
How it works
| AdamW (8-bit) | TACO (Ours) | |
|---|---|---|
| Persistent Optimizer State Memory | 27.7 GB | 0.16 GB (省 174 倍) |
| Peak Training Memory | 80.6 GB | 27.5 GB (省 2.9 倍) |
| Max Model Tunable on Single H100 | ~13B 參數模型 | 30B - 32B 參數模型 |
| Accuracy & Converging Speed | 基準 (Baseline) | 與 AdamW 相當 (Comparable) |
Why it matters
Traditional LLM tuning relies heavily on PEFT methods like LoRA due to AdamW's heavy memory footprints. TACO proves that optimizer states can be shrunk to near-zero via elegant steepest-descent geometry without losing full-parameter update capacity. This democratizes large-scale fine-tuning, allowing academic labs and startups to fine-tune 30B-class models on a single GPU workstation instead of relying on expensive multi-GPU clusters.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Full-parameter fine-tuning of 30B-32B open-source LLMs (e.g., Llama) on a single 80GB H100 GPU.
- 2Fine-tuning models pre-trained with AdamW in VRAM-constrained environments without encountering geometric mismatch issues.
Limitations & caveats
- Specifically designed for 2D weight matrices; may require alternative handling for non-2D weights such as biases or 1D embeddings.
- As a newly proposed ultra-sparse optimizer, its long-term stability and hyperparameter sensitivity on 70B+ scale models remain to be validated.
Related
Hierarchical Continuous Diffusion Language Models: Coupling Discrete Tokens with Continuous Latents
層級連續擴散語言模型:結合離散 Token 與連續潛在軌跡的全新生成架構
HC-DLM couples discrete token generation with a continuous latent trajectory in a unified denoising process, overcoming key bottlenecks of prior diffusion language models in reasoning and generation tasks.

Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Trillion-Parameter MoEs
Olmo-core 3 登場:開源兆級參數 MoE 模型訓練基礎架構
Olmo-core 3 is an open-source training framework designed to scale Mixture-of-Experts (MoE) models to the trillion-parameter range while significantly reducing communication bottlenecks.
SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability
SCAPO:藉由半事實穩定性最佳化 RLVR 的 Token 級信用分配
SCAPO is a novel variant of GRPO that incorporates semifactual stability into token-level credit assignment, significantly improving LLM reasoning accuracy and out-of-distribution generalization.