Base Models Can Reason: Unlocking Latent Performance with Strategic Starting Tokens
基礎模型也能推理:啟動關鍵「開頭 token」釋放隱藏實力
This study demonstrates that base language models possess latent reasoning capabilities that can be unlocked by forcing specific starting token cues (such as ".\n\nOkay" or "Alright,"). For instance, Olmo-3-7B's MATH-500 accuracy jumps from 42% to 78% when cued. The authors show that RL primarily increases the likelihood of generating these existing cues. Through causal data interventions, they successfully turned arbitrary words like 'chicken' into reasoning triggers, proving that these cues map to specific document types in the pre-training data.
Key points
Cues Unlock Latent Reasoning
Fixing specific starting tokens like ".\n\nOkay" elevates base model performance on math and coding tasks to be competitive with RL-trained models.
Rethinking Reinforcement Learning
RL primarily acts by making these effective reasoning-eliciting starting cues more likely to be generated by the model.
Causal Data Interventions
Researchers successfully trained arbitrary words like "chicken" or nonsense instructions to act as reasoning cues through data interventions.
Linking to Pre-training Data
Distinct token cues induce hidden states that correlate with specific document types from the pre-training dataset.
Why it matters
This research reshapes our understanding of how reasoning emerges in LLMs, showing that it is largely latent in pre-training data rather than purely injected via RL. By understanding token-level triggers, developers can elicit high-quality reasoning and control model safety directly from base models, reducing reliance on expensive RL training pipelines.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Zero-shot base model steering: Boosting math and coding performance on base models without expensive RL fine-tuning.
- 2Lightweight safety steering: Utilizing targeted starting tokens to elicit refusal or compliance behaviors.
Limitations & caveats
- Requires manual search or testing to find the optimal starting tokens for a given model (e.g., 'Okay' for Olmo vs. 'Alright' for Qwen).
- Evaluated primarily on specific domains like math, coding, and safety alignment; generalizability to all complex tasks is yet to be fully mapped.
Related

Falcon-Emirati-7B: Bridging the Gap in Emirati Arabic Dialect and Culture
解鎖阿聯酋方言與文化:專為在地語境打造的 Falcon-Emirati-7B 模型
Falcon-Emirati-7B is a 7B parameter model specialized in Emirati Arabic, capturing local dialect, Nabati poetry, and cultural nuances where generic models fail.
LESSER: High-Efficiency Post-Training Data Selection Using Output-Layer Gradients
LESSER:僅用輸出層梯度,實現高達 9.7 倍加速的 LLM 訓練後資料篩選
Researchers introduce LESSER, a method that uses output-layer gradients instead of full-parameter gradients to select post-training data, slashing compute costs by up to 9.7x while maintaining downstream performance.
TACO Optimizer: Unleashing Full-Parameter 32B LLM Fine-Tuning on a Single GPU
TACO 最佳化器:將 32B 大模型全參數微調帶入單張 GPU 的極簡幾何學
TACO is an ultra-low-memory optimizer that reduces persistent optimizer states by 174x, allowing full-parameter fine-tuning of 32B models on a single 80GB GPU.