arXivLLM
LESSER: High-Efficiency Post-Training Data Selection Using Output-Layer Gradients
LESSER:僅用輸出層梯度,實現高達 9.7 倍加速的 LLM 訓練後資料篩選
Researchers introduce LESSER, a method that uses output-layer gradients instead of full-parameter gradients to select post-training data, slashing compute costs by up to 9.7x while maintaining downstream performance.
2 min read