LESSER: High-Efficiency Post-Training Data Selection Using Output-Layer Gradients
LESSER:僅用輸出層梯度,實現高達 9.7 倍加速的 LLM 訓練後資料篩選
Selecting the right post-training data is crucial for LLM alignment, but traditional gradient-based selection is computationally prohibitive because it requires expensive backward passes for full gradients across all candidates. LESSER resolves this bottleneck by utilizing only output-layer gradients, which can be extracted at the cost of a cheap forward pass. It reduces feature-extraction compute costs by 9.7x for SFT and 3.0x for RL benchmarks while matching the downstream task performance of full-gradient methods.
Key points
No full backpropagation needed
LESSER demonstrates that output-layer gradients are sufficient for data selection, bypassing the costly backward passes required for full-parameter gradients.
Massive computational savings
It reduces feature-extraction FLOP costs by 9.7x for SFT and 3.0x for RL benchmarks.
Matches full-gradient performance
Despite the reduced compute, the selected datasets track full-gradient performance closely on downstream tasks.
Empirical batch alignment
Empirically, even if output-layer and full gradients rank individual samples differently, they successfully select batches with aligned gradients.
Why it matters
As LLMs scale, selecting high-quality data for customized alignment (SFT/RLHF) has become a massive bottleneck. LESSER challenges the assumption that accurate gradient-based selection requires full-model backpropagation. By showing that output-layer gradients from a cheap forward pass are sufficient, it lowers the computational barrier for customized post-training, enabling resource-constrained teams to curate optimal datasets at a fraction of the cost.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Efficient data selection for Supervised Fine-Tuning (SFT)
- 2Low-cost dataset curation for Reinforcement Learning (RL) alignment
Limitations & caveats
- It may rank individual samples differently than full-gradient methods, relying on batch-level alignment to achieve its efficacy.
- As a drop-in wrapper, its overall performance remains dependent on the underlying selection framework it integrates with.
Related
TACO Optimizer: Unleashing Full-Parameter 32B LLM Fine-Tuning on a Single GPU
TACO 最佳化器:將 32B 大模型全參數微調帶入單張 GPU 的極簡幾何學
TACO is an ultra-low-memory optimizer that reduces persistent optimizer states by 174x, allowing full-parameter fine-tuning of 32B models on a single 80GB GPU.
Hierarchical Continuous Diffusion Language Models: Coupling Discrete Tokens with Continuous Latents
層級連續擴散語言模型:結合離散 Token 與連續潛在軌跡的全新生成架構
HC-DLM couples discrete token generation with a continuous latent trajectory in a unified denoising process, overcoming key bottlenecks of prior diffusion language models in reasoning and generation tasks.

Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Trillion-Parameter MoEs
Olmo-core 3 登場:開源兆級參數 MoE 模型訓練基礎架構
Olmo-core 3 is an open-source training framework designed to scale Mixture-of-Experts (MoE) models to the trillion-parameter range while significantly reducing communication bottlenecks.