Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Pivot-SD:針對遮罩擴散語言模型的高效自我蒸餾框架
Masked diffusion language models (dLMs) struggle with credit assignment during post-training, as only a few critical denoising commitments (pivots) determine the final output's success. Pivot-SD solves this by using an information-gain metric to isolate these high-impact pivots. Pivots from successful trajectories are trained using cross-entropy, while those from failed trajectories are optimized using targeted unlikelihood, leaving the rest of the sequence untouched. Using only 200 training questions with four rollouts each, Pivot-SD outperforms full-sequence SFT and budget-matched RL baselines on math and code tasks.
Key points
Surgical Credit Assignment
Addresses the credit-assignment challenge in dLMs by focusing computational updates exclusively on high-impact tokens rather than the entire sequence.
Information-Gain Pivot Selection
Identifies pivotal decision points by measuring which token commitments yield the largest reduction in uncertainty over remaining masked tokens.
Asymmetric Optimization
Applies cross-entropy to successful pivots and targeted unlikelihood to failed pivots, shielding correct segments of failed runs from negative gradients.
Extreme Sample Efficiency
Achieves superior results over full-sequence SFT and budget-matched diffusion RL baselines using only 200 questions and 4 rollouts.
How it works
Why it matters
While parallel-generation-capable dLMs hold immense potential, post-training alignment has been complex and computationally expensive. Pivot-SD demonstrates that precise, localized updates to key decision points can significantly enhance reasoning capabilities. This provides an elegant, highly resource-efficient alignment framework for non-autoregressive architectures, making them competitive with traditional autoregressive models on complex tasks.
Who it affects
- AI Researcher
- AI Developer
How to use it
- 1Enhancing the mathematical reasoning and coding capabilities of parallel masked dLMs like LLaDA under strict compute constraints.
- 2Applying offline self-distillation for rapid post-training alignment and calibration of non-autoregressive LLMs.
Limitations & caveats
- Evaluation is primarily limited to math and coding tasks; performance on open-ended generation and creative writing remains to be explored.
- Relies on explicit information-gain calculations, which introduces processing overhead during the offline training data preparation phase.
Related
Less Decoder is More Encoder: Extracting Robust 3D Geometric Representations via Novel View Synthesis
減少解碼器反而增強編碼器:從新視角合成中提煉強大三維幾何表徵
This paper reveals how expressive decoders dilute geometric learning in Novel View Synthesis, and proposes SNAP—a self-supervised framework that restricts decoders to force encoders to learn robust 3D representations.
4DCodeBench: Benchmarking AI Agents on 4D Inverse Graphics and Dynamic Scene Code Generation
4DCodeBench:評估 AI Agent 動態場景 4D 反向圖形學與程式碼生成能力的全新基準
4DCodeBench is a new benchmark designed to evaluate AI agents' ability to reconstruct 4D dynamic scenes from videos by generating executable graphics and physics code.
What Should World Models Forget? Stratified Retention for Continual Adaptation
世界模型該遺忘什麼?以「分層保留」實現持續適應環境的能力
This paper argues that world models must not avoid all forgetting, proposing 'stratified retention' to distinguish permanent physical laws from dynamic, environment-specific facts that require timely revision.