SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability
SCAPO:藉由半事實穩定性最佳化 RLVR 的 Token 級信用分配
In reinforcement learning with verifiable rewards (RLVR), standard GRPO assigns a uniform outcome-derived advantage to all tokens, which can mistakenly reinforce spurious prompt features. To address this, researchers developed SCAPO (Semifactual Credit-Augmented Policy Optimization). It utilizes semifactual prompt interventions to calculate token probability drift. By reducing the credit advantage of highly unstable tokens during early training, SCAPO dramatically reduces prompt sensitivity. Evaluation on Qwen3 base models shows significant improvements in AIME benchmarks and out-of-distribution mathematical reasoning tasks.
Key points
Flaw in GRPO's uniform credit assignment
Standard GRPO assigns the exact same outcome advantage to all tokens, risking the reinforcement of spurious prompt-dependent features.
Semifactual interventions & drift measurement
SCAPO measures token probability drift for fixed responses under semifactual prompt interventions that preserve the core problem.
Discounting unstable token advantages
Using normalized stability scores, SCAPO penalizes the advantage values of highly unstable tokens without granting extra credit for stability alone.
Boosted reasoning & OOD generalization
On Qwen3-4B and 1.7B, SCAPO improved AIME accuracy by 5.63 and 4.17 percentage points, achieving top results on OOD benchmarks.
How it works
Why it matters
When training LLMs via RL, models often learn the right answer through spurious correlation. SCAPO addresses this with a causally inspired, fine-grained credit assignment mechanism. By tackling prompt sensitivity directly at the token level, it builds robust reasoning models that generalize far better to out-of-distribution math tasks without relying on brittle prompt configurations.
Who it affects
- AI Researcher
- AI Developer
How to use it
- 1Reinforcement learning training and alignment for reasoning LLMs
- 2Mitigating model sensitivity toward prompt formats and templates
Limitations & caveats
- Requires designing semifactual prompt variations, potentially increasing upfront data curation and computational overhead.
- The core mechanism is primarily optimized for reasoning tasks with verifiable rewards, such as mathematics.
Related
STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸
STEPQuant is a spatial-temporal post-training quantization framework for Delta-rule recurrent states, achieving over 5x state compression and reducing serving memory by up to 68.7% with minimal accuracy loss.
LeapQuant: Near-Lossless 8-Bit Recurrent State Quantization for Linear Attention LLMs
LeapQuant:突破線性注意力瓶頸,實現近乎無損的 8-bit 遞迴狀態量化
LeapQuant is a training-free quantization method that achieves near-lossless 8-bit recurrent state quantization for linear attention, delivering up to 1.47x end-to-end inference speedup while preserving FP32 accuracy.
AdviSD: Steering Frontier LLMs via Targeted Multi-Turn Self-Distillation
微型顧問的逆襲:AdviSD 透過目標多輪自我蒸餾精準引導前沿大模型
AdviSD enables a small trainable advisor to steer frozen frontier LLMs using natural-language advice, leveraging selective self-distillation to filter and learn only from high-impact corrections.