Aivora
arXivAI ResearchAdvanced

Planning to Learn: Bridging Cross-Entropy and Policy Gradient with Horizon Loss

規劃學習:以 Horizon Loss 橋接交叉熵與策略梯度的全新損失函數

2 min read
Planning to Learn: Bridging Cross-Entropy and Policy Gradient with Horizon Loss
The 30-second version

Policy-gradient methods power modern LLM post-training, yet in simple classification, exact policy gradient (EPG) surprisingly underperforms against cross-entropy (CE). The author identifies that EPG is myopic, optimizing only for immediate rewards. Conversely, CE acts as 'patient accuracy' over an infinite horizon. To resolve this, the paper introduces 'horizon loss'—a simple one-line code modification that truncates the horizon to match remaining training steps. Tested on MNIST and ImageNet (ResNet, ViT), horizon loss significantly improves top-1 accuracy, with performance gains scaling alongside label noise.

Key points

01

The Myopia of Policy Gradients

Exact policy gradient (EPG) only values immediate updates, ignoring how the current step positions the model for subsequent learning.

02

Cross-Entropy as Patient Accuracy

Cross-entropy (CE) represents the total error over an infinite horizon, making it inherently more long-sighted than standard EPG.

03

A Simple One-Line Solution

By truncating the learning horizon to match the remaining training steps, this one-line loss modification blends the strengths of CE and EPG.

04

Robustness Against Label Noise

Evaluated on MNIST and ImageNet (ResNet/ViT), horizon loss outperforms CE, with performance gains scaling alongside label noise.

How it works

Comparison of Cross-Entropy, EPG, and Horizon Loss
Cross-Entropy (CE)Exact Policy Gradient (EPG)Horizon Loss
Learning Horizon無限地平線 (Infinite)零地平線 (Zero-horizon)動態遞減 (Decaying with training)
Optimization Style有耐心的長線學習近視、僅關注當下回報前期重長線,後期精準收斂
Noise Tolerance一般 (Moderate)較差 (Poor)優異 (Superior)

Why it matters

As policy gradients are crucial for LLM RL post-training, understanding their inherent myopia is highly impactful. This paper offers an elegant, zero-overhead mathematical fix (horizon loss) that improves training efficiency and noise tolerance, building a key conceptual bridge between supervised loss functions and RL planning.

Who it affects

  • AI Developer
  • AI Researcher

How to use it

  1. 1Deep learning image classification (e.g., training ResNet and ViT on ImageNet)
  2. 2Training on real-world datasets with high rates of label noise
  3. 3Optimizing policy-gradient-based RL and LLM post-training alignment

Limitations & caveats

  • Empirical validation is currently focused on classification tasks, and has not yet been fully demonstrated in large-scale LLM RL environments.
  • Deploying horizon loss requires tracking or defining the remaining training horizon, which may be complex in open-ended training scenarios.

Related

What Should World Models Forget? Stratified Retention for Continual Adaptation
arXivAI Research

What Should World Models Forget? Stratified Retention for Continual Adaptation

世界模型該遺忘什麼?以「分層保留」實現持續適應環境的能力

This paper argues that world models must not avoid all forgetting, proposing 'stratified retention' to distinguish permanent physical laws from dynamic, environment-specific facts that require timely revision.

2 min read