Strategically Diverse Sampling: Why Approach Diversity Beats Model Scale in Self-Training
自我訓練新突破:策略多樣性取樣,如何讓小模型靠「多元思路」擊敗大模型蒸餾?
Traditional LLM self-training relies on IID sampling filtered by correctness, which biases models toward already-favored strategies. This paper proposes "strategically diverse sampling" using GROOT (a hierarchical approach tree) and Verbalized Sampling (VS) to generate varied reasoning paths. Strikingly, self-training on diverse but incorrect traces from a Qwen3-4B model outperformed IID distillation from a 235B teacher, proving that diversity of thought matters more than absolute correctness or teacher size.
Key points
Limitations of IID Sampling
Standard self-training relies on repeated IID sampling and filtering for correctness, which overrepresents strategies the model already prefers.
Strategic Diversity Sampling
Focuses on substantive variations in problem-solving approaches. Introduces GROOT (hierarchical approach tree) and VS (verbalized sampling) to extract diverse paths.
Diversity Beats Teacher Scale
Training on strategically diverse but incorrect traces from Qwen3-4B unexpectedly outperforms IID distillation from a massive 235B teacher model.
Boost for RL & Test-Time Scaling
Provides a strong initialization baseline that significantly enhances subsequent reinforcement learning and test-time scaling performance.
How it works
| 傳統 IID 取樣 | 策略多樣性取樣 | |
|---|---|---|
| Core Driver | 答案正確性與教師規模 | 解題思路的實質多樣性 |
| Sampling | 獨立同分布重複取樣 | 層級策略樹或口語化取樣 |
| Small Model Results | 高度依賴大模型蒸餾 | 錯誤但多樣的資料仍能勝過蒸餾 |
Why it matters
This research challenges the long-held assumption that self-training data must be strictly correct or sourced from massive teacher models. By demonstrating that strategic diversity is more critical than correctness or model scale, it opens up a cost-effective path for smaller teams. Developers can now leverage smaller models to generate diverse reasoning paths, achieving competitive or superior results compared to distilling from expensive 235B models.
Who it affects
- AI Developer
- AI Researcher
How to use it
- 1Generating multi-approach solutions for self-training in competitive programming.
- 2Preparing high-quality, diverse initialization datasets for RL and Chain-of-Thought reasoning.
Limitations & caveats
- Currently evaluated only on competitive programming and Next-Chapter Prediction; generalizability to other domains requires further validation.
- Constructing hierarchical decision trees (GROOT) and verbalized sampling prompts may introduce additional compute and design overhead.
Related
FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
FurE:免用動物毛髮資料集,實現 10 倍加速的 3D 動物毛髮重建技術
FurE is an efficient 3D animal fur reconstruction method that leverages a human-hair trained PCA decoder and Gaussian Frosting to achieve 10x faster, highly detailed, and editable groom reconstruction without animal datasets.
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率
Researchers found that training reasoning models to predict their own confidence at intermediate steps naturally reduces generated tokens by up to 25% at matched accuracy, without explicitly optimizing for length or stopping.
First-Order Stationarity of Reverse Diffusions: Bridging Optimization and Sampling
逆向擴散的一階駐點性:連結優化與生成採樣的數學機制
This research establishes a first-order optimization theory for diffusion models, proving that SDE-based reverse Langevin diffusions contract Fisher divergences exponentially under strongly convex noising.