Aivora
arXivLLMAdvanced

LESSER: High-Efficiency Post-Training Data Selection Using Output-Layer Gradients

LESSER:僅用輸出層梯度,實現高達 9.7 倍加速的 LLM 訓練後資料篩選

2 min read
LESSER: High-Efficiency Post-Training Data Selection Using Output-Layer Gradients
The 30-second version

Selecting the right post-training data is crucial for LLM alignment, but traditional gradient-based selection is computationally prohibitive because it requires expensive backward passes for full gradients across all candidates. LESSER resolves this bottleneck by utilizing only output-layer gradients, which can be extracted at the cost of a cheap forward pass. It reduces feature-extraction compute costs by 9.7x for SFT and 3.0x for RL benchmarks while matching the downstream task performance of full-gradient methods.

Key points

01

No full backpropagation needed

LESSER demonstrates that output-layer gradients are sufficient for data selection, bypassing the costly backward passes required for full-parameter gradients.

02

Massive computational savings

It reduces feature-extraction FLOP costs by 9.7x for SFT and 3.0x for RL benchmarks.

03

Matches full-gradient performance

Despite the reduced compute, the selected datasets track full-gradient performance closely on downstream tasks.

04

Empirical batch alignment

Empirically, even if output-layer and full gradients rank individual samples differently, they successfully select batches with aligned gradients.

Why it matters

As LLMs scale, selecting high-quality data for customized alignment (SFT/RLHF) has become a massive bottleneck. LESSER challenges the assumption that accurate gradient-based selection requires full-model backpropagation. By showing that output-layer gradients from a cheap forward pass are sufficient, it lowers the computational barrier for customized post-training, enabling resource-constrained teams to curate optimal datasets at a fraction of the cost.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Efficient data selection for Supervised Fine-Tuning (SFT)
  2. 2Low-cost dataset curation for Reinforcement Learning (RL) alignment

Limitations & caveats

  • It may rank individual samples differently than full-gradient methods, relying on batch-level alignment to achieve its efficacy.
  • As a drop-in wrapper, its overall performance remains dependent on the underlying selection framework it integrates with.

Related