Aivora
arXivLLMAdvanced

TACO Optimizer: Unleashing Full-Parameter 32B LLM Fine-Tuning on a Single GPU

TACO 最佳化器:將 32B 大模型全參數微調帶入單張 GPU 的極簡幾何學

2 min read
TACO Optimizer: Unleashing Full-Parameter 32B LLM Fine-Tuning on a Single GPU
The 30-second version

Full-parameter LLM fine-tuning is heavily bottlenecked by the massive state memory of optimizers like AdamW. This paper introduces TACO, which computes updates by selecting only the sign of the absolute-maximum gradient entry in each column of the weight matrix. On OPT-13B, TACO reduces persistent optimizer states by 174x (from 27.7 GB to 0.16 GB) and peak training memory by 2.9x compared to 8-bit AdamW. Maintaining comparable accuracy and runtime, TACO enables full-parameter fine-tuning of 30-32B models on a single 80GB H100 GPU.

Key points

01

Negligible Optimizer State

Reduces persistent optimizer state from 27.7 GB to 0.16 GB for OPT-13B, representing a 174x memory saving compared to 8-bit AdamW.

02

Column-wise One-sparse Geometry

Computes exact steepest descent by extracting only the sign of the largest magnitude entry in each column of 2D weight matrices.

03

Single-GPU Scaling Boost

Enables full-parameter fine-tuning of 30-32B parameter models across multiple model families on a single 80 GB H100 GPU.

04

Comparable Fine-tuning Accuracy

Overcomes geometric discrepancy issues of prior approaches like Muon, ensuring convergence quality and accuracy on par with AdamW.

How it works

OPT-13B Fine-Tuning Metrics: AdamW8bit vs TACO
AdamW (8-bit)TACO (Ours)
Persistent Optimizer State Memory27.7 GB0.16 GB (省 174 倍)
Peak Training Memory80.6 GB27.5 GB (省 2.9 倍)
Max Model Tunable on Single H100~13B 參數模型30B - 32B 參數模型
Accuracy & Converging Speed基準 (Baseline)與 AdamW 相當 (Comparable)

Why it matters

Traditional LLM tuning relies heavily on PEFT methods like LoRA due to AdamW's heavy memory footprints. TACO proves that optimizer states can be shrunk to near-zero via elegant steepest-descent geometry without losing full-parameter update capacity. This democratizes large-scale fine-tuning, allowing academic labs and startups to fine-tune 30B-class models on a single GPU workstation instead of relying on expensive multi-GPU clusters.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Full-parameter fine-tuning of 30B-32B open-source LLMs (e.g., Llama) on a single 80GB H100 GPU.
  2. 2Fine-tuning models pre-trained with AdamW in VRAM-constrained environments without encountering geometric mismatch issues.

Limitations & caveats

  • Specifically designed for 2D weight matrices; may require alternative handling for non-2D weights such as biases or 1D embeddings.
  • As a newly proposed ultra-sparse optimizer, its long-term stability and hyperparameter sensitivity on 70B+ scale models remain to be validated.

Related

SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability
arXivLLM

SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability

SCAPO:藉由半事實穩定性最佳化 RLVR 的 Token 級信用分配

SCAPO is a novel variant of GRPO that incorporates semifactual stability into token-level credit assignment, significantly improving LLM reasoning accuracy and out-of-distribution generalization.

2 min read