Aivora
arXivAI ResearchIntermediate

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率

2 min read
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
The 30-second version

To reduce the length of reasoning chains, current approaches rely on inference-time early stopping or reinforcement learning with length penalties. This paper proposes a self-supervised confidence training method using only 600 training problems. The model is fine-tuned solely to predict its confidence at intermediate steps, with no training objective for length or efficiency. Remarkably, during standard inference without early stopping, the fine-tuned models naturally generate up to 25% fewer tokens while maintaining accuracy. This efficiency gain is consistent across Gemma, Qwen, Nemotron, and GPT-OSS models.

Key points

01

Metacognitive Signal Training

Fine-tunes reasoning models using a self-supervised procedure to predict their confidence in the final answer at intermediate steps.

02

Zero Efficiency Objectives

The training loss contains no parameters or goals for length, efficiency, or early-stopping mechanisms during inference.

03

Up to 25% Token Reduction

Reduces generated tokens by up to 25% at matched accuracy across multiple reasoning benchmarks and base models.

04

Preserved Reasoning Logic

Analysis shows that confidence training largely preserves the base models' high-level reasoning composition instead of suppressing specific behaviors.

Why it matters

This research reveals that efficient reasoning can naturally emerge as a downstream consequence of learning metacognitive self-assessment. Developers can significantly lower inference costs for reasoning models without complex RL pipelines or custom early-stopping logic, requiring only a lightweight dataset of 600 problems for fine-tuning.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Reducing token costs and latency for Chain-of-Thought reasoning models in enterprise production.
  2. 2Optimizing resource consumption of math, science, and coding assistants without degrading quality.

Limitations & caveats

  • The approach is primarily evaluated on math, science, and coding benchmarks; its efficacy on general chat or creative tasks remains untested.
  • While effective across multiple base models, the optimal setup for confidence formulation might still vary slightly by architecture.

Related

New LoRA Skills Should Read but Never Write: READ Solves Adapter Fusion Interference
arXivAI Research

New LoRA Skills Should Read but Never Write: READ Solves Adapter Fusion Interference

新增 LoRA 技能唯讀不寫:READ 解決多配接器融合干擾

This paper introduces READ, a method that resolves interference when merging multiple LoRA adapters by enforcing "read-only" one-way coupling and canonical factorization, preserving old skills while adding new ones with zero extra inference cost.

2 min read