Aivora
arXivAI ResearchAdvanced

Decoupling Direction and Step-Size: ZFO Framework for LLM Fine-Tuning

結合零階與一階優化:ZFO 框架為大語言模型微調尋找最佳步長

2 min read
Decoupling Direction and Step-Size: ZFO Framework for LLM Fine-Tuning
The 30-second version

Selecting the optimal step-size in LLM fine-tuning is challenging, as bad choices lead to slow convergence or training instability. The proposed Zero-and-First-Order (ZFO) framework solves this by decoupling direction and step-size selection. ZFO utilizes standard first-order gradients to find the optimization direction, then performs only two zeroth-order objective evaluations along this 1D path to construct a local curvature model. This allows it to compute an adaptive, curvature-aware step size at a fraction of the cost of a full line search, outperforming standard fixed-step first-order baselines.

Key points

01

Decoupled Optimization

Uses standard first-order optimizers to choose the search direction, and zeroth-order evaluations to decide how far to move.

02

Low Computational Overhead

Requires only current gradient information and two additional objective evaluations, making it much cheaper than a full line search.

03

Curvature-Aware Adaptation

Builds a local model along the proposed 1D direction using finite-difference curvature estimates to choose near-optimal steps.

04

Theoretical Guarantees

Proves that shared-sample evaluations produce reliable curvature estimates and that ZFO converges to a neighborhood of a stationary point.

How it works

ZFO Step-Size Adaptive Optimization Process
Along 1D subspaceEstimate finite-differenceSearch bounded intervalApply step sizeDetermine Direction2x Zeroth-Order EvalBuild Curvature ModelSelect Near-OptimalStepUpdate Parameters

Why it matters

Standard LLM fine-tuning relies heavily on manual learning rate tuning and schedules, which can waste immense compute if misconfigured. ZFO introduces an automatic, lightweight step-size adaptation mechanism. By requiring only two extra forward passes to compute near-optimal steps, it significantly improves training stability and final model performance, offering a practical optimization upgrade for large-scale training.

Who it affects

  • AI Developer
  • AI Researcher

How to use it

  1. 1Fine-tuning LLMs where manual learning rate scheduling is unstable or highly resource-intensive.
  2. 2Adaptive optimization in resource-constrained settings, serving as a low-cost alternative to full line search.

Limitations & caveats

  • Requires two additional objective evaluations (forward passes) per step, which introduces minor computational latency.
  • The preferred local model configuration and overall performance gains depend heavily on the specific dataset and optimization objective.

Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
Hugging FaceAI Research

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude

艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型

Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.

2 min read