Decoupling Direction and Step-Size: ZFO Framework for LLM Fine-Tuning
結合零階與一階優化:ZFO 框架為大語言模型微調尋找最佳步長
Selecting the optimal step-size in LLM fine-tuning is challenging, as bad choices lead to slow convergence or training instability. The proposed Zero-and-First-Order (ZFO) framework solves this by decoupling direction and step-size selection. ZFO utilizes standard first-order gradients to find the optimization direction, then performs only two zeroth-order objective evaluations along this 1D path to construct a local curvature model. This allows it to compute an adaptive, curvature-aware step size at a fraction of the cost of a full line search, outperforming standard fixed-step first-order baselines.
Key points
Decoupled Optimization
Uses standard first-order optimizers to choose the search direction, and zeroth-order evaluations to decide how far to move.
Low Computational Overhead
Requires only current gradient information and two additional objective evaluations, making it much cheaper than a full line search.
Curvature-Aware Adaptation
Builds a local model along the proposed 1D direction using finite-difference curvature estimates to choose near-optimal steps.
Theoretical Guarantees
Proves that shared-sample evaluations produce reliable curvature estimates and that ZFO converges to a neighborhood of a stationary point.
How it works
Why it matters
Standard LLM fine-tuning relies heavily on manual learning rate tuning and schedules, which can waste immense compute if misconfigured. ZFO introduces an automatic, lightweight step-size adaptation mechanism. By requiring only two extra forward passes to compute near-optimal steps, it significantly improves training stability and final model performance, offering a practical optimization upgrade for large-scale training.
Who it affects
- AI Developer
- AI Researcher
How to use it
- 1Fine-tuning LLMs where manual learning rate scheduling is unstable or highly resource-intensive.
- 2Adaptive optimization in resource-constrained settings, serving as a low-cost alternative to full line search.
Limitations & caveats
- Requires two additional objective evaluations (forward passes) per step, which introduces minor computational latency.
- The preferred local model configuration and overall performance gains depend heavily on the specific dataset and optimization objective.
Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型
Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.
GALA: Distilling 3D Gaussian Avatars into Linear Blendshapes for Real-Time Animation
GALA:用線性混合變形蒸餾技術實現 3D Gaussian 虛擬化身即時動畫
GALA distills complex neural decoding of 3D Gaussian avatars into lightweight linear blendshapes, reducing CPU animation costs by up to 1000x and enabling 60fps real-time performance on mobile devices.
ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers
ScholarCatalyst:評估 AI 是否擁有「科學家直覺」的學術文獻檢索基準
ScholarCatalyst is a novel benchmark featuring annotations from 184 lead authors to evaluate whether AI can retrieve key inspiring papers from past literature based only on an initial research question.