Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率
To reduce the length of reasoning chains, current approaches rely on inference-time early stopping or reinforcement learning with length penalties. This paper proposes a self-supervised confidence training method using only 600 training problems. The model is fine-tuned solely to predict its confidence at intermediate steps, with no training objective for length or efficiency. Remarkably, during standard inference without early stopping, the fine-tuned models naturally generate up to 25% fewer tokens while maintaining accuracy. This efficiency gain is consistent across Gemma, Qwen, Nemotron, and GPT-OSS models.
Key points
Metacognitive Signal Training
Fine-tunes reasoning models using a self-supervised procedure to predict their confidence in the final answer at intermediate steps.
Zero Efficiency Objectives
The training loss contains no parameters or goals for length, efficiency, or early-stopping mechanisms during inference.
Up to 25% Token Reduction
Reduces generated tokens by up to 25% at matched accuracy across multiple reasoning benchmarks and base models.
Preserved Reasoning Logic
Analysis shows that confidence training largely preserves the base models' high-level reasoning composition instead of suppressing specific behaviors.
Why it matters
This research reveals that efficient reasoning can naturally emerge as a downstream consequence of learning metacognitive self-assessment. Developers can significantly lower inference costs for reasoning models without complex RL pipelines or custom early-stopping logic, requiring only a lightweight dataset of 600 problems for fine-tuning.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Reducing token costs and latency for Chain-of-Thought reasoning models in enterprise production.
- 2Optimizing resource consumption of math, science, and coding assistants without degrading quality.
Limitations & caveats
- The approach is primarily evaluated on math, science, and coding benchmarks; its efficacy on general chat or creative tasks remains untested.
- While effective across multiple base models, the optimal setup for confidence formulation might still vary slightly by architecture.
Related
First-Order Stationarity of Reverse Diffusions: Bridging Optimization and Sampling
逆向擴散的一階駐點性:連結優化與生成採樣的數學機制
This research establishes a first-order optimization theory for diffusion models, proving that SDE-based reverse Langevin diffusions contract Fisher divergences exponentially under strongly convex noising.
Statistical Attribute Alignment for Black-Box Generative AI via Output Post-Processing
黑盒生成式 AI 的統計屬性對齊:透過輸出後處理實現公平與多樣性
This paper introduces post-processing algorithms to align the attribute distribution of black-box generative AI outputs with user-specified targets using a mathematically minimized number of queries.
New LoRA Skills Should Read but Never Write: READ Solves Adapter Fusion Interference
新增 LoRA 技能唯讀不寫:READ 解決多配接器融合干擾
This paper introduces READ, a method that resolves interference when merging multiple LoRA adapters by enforcing "read-only" one-way coupling and canonical factorization, preserving old skills while adding new ones with zero extra inference cost.