Aivora
arXivAI ResearchAdvanced

FERPO: Forward Entropy-Regularized Policy Optimization Without Action Gradients

FERPO:免除動作梯度的前向熵正規化策略最佳化,實現高效且穩定的強化學習

2 min read
FERPO: Forward Entropy-Regularized Policy Optimization Without Action Gradients
The 30-second version

Traditional continuous control RL relies on the action gradients of a learned critic, which are often unreliable. To resolve this, researchers introduced FERPO. Instead of differentiating the critic, FERPO derives an optimal target action distribution regularized by entropy and KL divergence. It then fits the actor using a forward-KL objective estimated via self-normalized importance sampling (SNIS). This avoids unstable gradient updates and encourages covering multiple high-value modes for better exploration, delivering superior sample efficiency and faster actor updates on MuJoCo and ManiSkill benchmarks.

Key points

01

Action-Gradient-Free Updates

Improves policies directly using critic values without computing action derivatives, avoiding unstable updates caused by inaccurate gradients.

02

Forward-KL for Better Exploration

Unlike reverse-KL which favors a subset of modes, the forward-KL objective encourages covering multiple high-value modes to promote exploration.

03

Stable SNIS Estimation

Limits target distribution deviation from the rollout policy using KL regularization, keeping SNIS importance weights well-behaved.

04

Faster Computation & Sample Efficiency

Demonstrates competitive performance and sample-efficiency gains, with faster actor updates than REPPO.

How it works

Comparison of FERPO and Traditional Continuous Control RL Methods
標準梯度型方法 (Gradient-Based)REPPO 演算法FERPO (本研究)
Critic Derivative Needed需要 (容易因微分誤差而不穩定)需要 (計算 Pathwise 梯度)不需要 (僅使用數值,更新極穩定)
Optimization Objective逆向 KL 或直接策略梯度逆向 KL 正規化前向 KL 正規化與 SNIS 估計
Exploration & Mode Coverage傾向於單一模態 (容易陷入局部最優)傾向於單一模態 (Mode-seeking)可覆蓋多個高價值模態 (Mode-covering)
Computation Speed標準較慢 (因 Pathwise 梯度計算開銷大)極快 (Actor 更新開銷大幅減少)

Why it matters

In continuous control domains like robotics, policy updates are often plagued by inaccurate critic gradients. FERPO provides an elegant gradient-free alternative. By utilizing a forward-KL objective, it naturally prevents mode collapse and enhances multimodal exploration. Combined with outstanding sample efficiency and low computational overhead on simulator benchmarks, FERPO paves the way for faster and more stable real-world robotic learning.

Who it affects

  • AI Researcher
  • AI Developer
  • Student & Learner

How to use it

  1. 1Robotic arm manipulation (e.g., high-precision grasping and assembly in ManiSkill)
  2. 2Legged locomotion control and motion planning for quadrupeds and bipeds in complex environments
  3. 3Sample-efficient online reinforcement learning simulations requiring rapid prototype iteration

Limitations & caveats

  • As an on-policy algorithm, its overall data efficiency may still be constrained in scenarios where off-policy learning is preferred.
  • If the target distribution deviates excessively from the rollout policy, SNIS may still suffer from high variance in importance weights.

Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
Hugging FaceAI Research

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude

艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型

Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.

2 min read