Aivora
arXivLLMIntermediate

Base Models Can Reason: Unlocking Latent Performance with Strategic Starting Tokens

基礎模型也能推理:啟動關鍵「開頭 token」釋放隱藏實力

2 min read
Base Models Can Reason: Unlocking Latent Performance with Strategic Starting Tokens
The 30-second version

This study demonstrates that base language models possess latent reasoning capabilities that can be unlocked by forcing specific starting token cues (such as ".\n\nOkay" or "Alright,"). For instance, Olmo-3-7B's MATH-500 accuracy jumps from 42% to 78% when cued. The authors show that RL primarily increases the likelihood of generating these existing cues. Through causal data interventions, they successfully turned arbitrary words like 'chicken' into reasoning triggers, proving that these cues map to specific document types in the pre-training data.

Key points

01

Cues Unlock Latent Reasoning

Fixing specific starting tokens like ".\n\nOkay" elevates base model performance on math and coding tasks to be competitive with RL-trained models.

02

Rethinking Reinforcement Learning

RL primarily acts by making these effective reasoning-eliciting starting cues more likely to be generated by the model.

03

Causal Data Interventions

Researchers successfully trained arbitrary words like "chicken" or nonsense instructions to act as reasoning cues through data interventions.

04

Linking to Pre-training Data

Distinct token cues induce hidden states that correlate with specific document types from the pre-training dataset.

Why it matters

This research reshapes our understanding of how reasoning emerges in LLMs, showing that it is largely latent in pre-training data rather than purely injected via RL. By understanding token-level triggers, developers can elicit high-quality reasoning and control model safety directly from base models, reducing reliance on expensive RL training pipelines.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Zero-shot base model steering: Boosting math and coding performance on base models without expensive RL fine-tuning.
  2. 2Lightweight safety steering: Utilizing targeted starting tokens to elicit refusal or compliance behaviors.

Limitations & caveats

  • Requires manual search or testing to find the optimal starting tokens for a given model (e.g., 'Okay' for Olmo vs. 'Alright' for Qwen).
  • Evaluated primarily on specific domains like math, coding, and safety alignment; generalizability to all complex tasks is yet to be fully mapped.

Related

LESSER: High-Efficiency Post-Training Data Selection Using Output-Layer Gradients
arXivLLM

LESSER: High-Efficiency Post-Training Data Selection Using Output-Layer Gradients

LESSER:僅用輸出層梯度,實現高達 9.7 倍加速的 LLM 訓練後資料篩選

Researchers introduce LESSER, a method that uses output-layer gradients instead of full-parameter gradients to select post-training data, slashing compute costs by up to 9.7x while maintaining downstream performance.

2 min read