Planning to Learn: Bridging Cross-Entropy and Policy Gradient with Horizon Loss
規劃學習:以 Horizon Loss 橋接交叉熵與策略梯度的全新損失函數
Policy-gradient methods power modern LLM post-training, yet in simple classification, exact policy gradient (EPG) surprisingly underperforms against cross-entropy (CE). The author identifies that EPG is myopic, optimizing only for immediate rewards. Conversely, CE acts as 'patient accuracy' over an infinite horizon. To resolve this, the paper introduces 'horizon loss'—a simple one-line code modification that truncates the horizon to match remaining training steps. Tested on MNIST and ImageNet (ResNet, ViT), horizon loss significantly improves top-1 accuracy, with performance gains scaling alongside label noise.
Key points
The Myopia of Policy Gradients
Exact policy gradient (EPG) only values immediate updates, ignoring how the current step positions the model for subsequent learning.
Cross-Entropy as Patient Accuracy
Cross-entropy (CE) represents the total error over an infinite horizon, making it inherently more long-sighted than standard EPG.
A Simple One-Line Solution
By truncating the learning horizon to match the remaining training steps, this one-line loss modification blends the strengths of CE and EPG.
Robustness Against Label Noise
Evaluated on MNIST and ImageNet (ResNet/ViT), horizon loss outperforms CE, with performance gains scaling alongside label noise.
How it works
| Cross-Entropy (CE) | Exact Policy Gradient (EPG) | Horizon Loss | |
|---|---|---|---|
| Learning Horizon | 無限地平線 (Infinite) | 零地平線 (Zero-horizon) | 動態遞減 (Decaying with training) |
| Optimization Style | 有耐心的長線學習 | 近視、僅關注當下回報 | 前期重長線,後期精準收斂 |
| Noise Tolerance | 一般 (Moderate) | 較差 (Poor) | 優異 (Superior) |
Why it matters
As policy gradients are crucial for LLM RL post-training, understanding their inherent myopia is highly impactful. This paper offers an elegant, zero-overhead mathematical fix (horizon loss) that improves training efficiency and noise tolerance, building a key conceptual bridge between supervised loss functions and RL planning.
Who it affects
- AI Developer
- AI Researcher
How to use it
- 1Deep learning image classification (e.g., training ResNet and ViT on ImageNet)
- 2Training on real-world datasets with high rates of label noise
- 3Optimizing policy-gradient-based RL and LLM post-training alignment
Limitations & caveats
- Empirical validation is currently focused on classification tasks, and has not yet been fully demonstrated in large-scale LLM RL environments.
- Deploying horizon loss requires tracking or defining the remaining training horizon, which may be complex in open-ended training scenarios.
Related
Less Decoder is More Encoder: Extracting Robust 3D Geometric Representations via Novel View Synthesis
減少解碼器反而增強編碼器:從新視角合成中提煉強大三維幾何表徵
This paper reveals how expressive decoders dilute geometric learning in Novel View Synthesis, and proposes SNAP—a self-supervised framework that restricts decoders to force encoders to learn robust 3D representations.
4DCodeBench: Benchmarking AI Agents on 4D Inverse Graphics and Dynamic Scene Code Generation
4DCodeBench:評估 AI Agent 動態場景 4D 反向圖形學與程式碼生成能力的全新基準
4DCodeBench is a new benchmark designed to evaluate AI agents' ability to reconstruct 4D dynamic scenes from videos by generating executable graphics and physics code.
What Should World Models Forget? Stratified Retention for Continual Adaptation
世界模型該遺忘什麼?以「分層保留」實現持續適應環境的能力
This paper argues that world models must not avoid all forgetting, proposing 'stratified retention' to distinguish permanent physical laws from dynamic, environment-specific facts that require timely revision.