FERPO: Forward Entropy-Regularized Policy Optimization Without Action Gradients
FERPO:免除動作梯度的前向熵正規化策略最佳化,實現高效且穩定的強化學習
Traditional continuous control RL relies on the action gradients of a learned critic, which are often unreliable. To resolve this, researchers introduced FERPO. Instead of differentiating the critic, FERPO derives an optimal target action distribution regularized by entropy and KL divergence. It then fits the actor using a forward-KL objective estimated via self-normalized importance sampling (SNIS). This avoids unstable gradient updates and encourages covering multiple high-value modes for better exploration, delivering superior sample efficiency and faster actor updates on MuJoCo and ManiSkill benchmarks.
Key points
Action-Gradient-Free Updates
Improves policies directly using critic values without computing action derivatives, avoiding unstable updates caused by inaccurate gradients.
Forward-KL for Better Exploration
Unlike reverse-KL which favors a subset of modes, the forward-KL objective encourages covering multiple high-value modes to promote exploration.
Stable SNIS Estimation
Limits target distribution deviation from the rollout policy using KL regularization, keeping SNIS importance weights well-behaved.
Faster Computation & Sample Efficiency
Demonstrates competitive performance and sample-efficiency gains, with faster actor updates than REPPO.
How it works
| 標準梯度型方法 (Gradient-Based) | REPPO 演算法 | FERPO (本研究) | |
|---|---|---|---|
| Critic Derivative Needed | 需要 (容易因微分誤差而不穩定) | 需要 (計算 Pathwise 梯度) | 不需要 (僅使用數值,更新極穩定) |
| Optimization Objective | 逆向 KL 或直接策略梯度 | 逆向 KL 正規化 | 前向 KL 正規化與 SNIS 估計 |
| Exploration & Mode Coverage | 傾向於單一模態 (容易陷入局部最優) | 傾向於單一模態 (Mode-seeking) | 可覆蓋多個高價值模態 (Mode-covering) |
| Computation Speed | 標準 | 較慢 (因 Pathwise 梯度計算開銷大) | 極快 (Actor 更新開銷大幅減少) |
Why it matters
In continuous control domains like robotics, policy updates are often plagued by inaccurate critic gradients. FERPO provides an elegant gradient-free alternative. By utilizing a forward-KL objective, it naturally prevents mode collapse and enhances multimodal exploration. Combined with outstanding sample efficiency and low computational overhead on simulator benchmarks, FERPO paves the way for faster and more stable real-world robotic learning.
Who it affects
- AI Researcher
- AI Developer
- Student & Learner
How to use it
- 1Robotic arm manipulation (e.g., high-precision grasping and assembly in ManiSkill)
- 2Legged locomotion control and motion planning for quadrupeds and bipeds in complex environments
- 3Sample-efficient online reinforcement learning simulations requiring rapid prototype iteration
Limitations & caveats
- As an on-policy algorithm, its overall data efficiency may still be constrained in scenarios where off-policy learning is preferred.
- If the target distribution deviates excessively from the rollout policy, SNIS may still suffer from high variance in importance weights.
Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型
Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.
GALA: Distilling 3D Gaussian Avatars into Linear Blendshapes for Real-Time Animation
GALA:用線性混合變形蒸餾技術實現 3D Gaussian 虛擬化身即時動畫
GALA distills complex neural decoding of 3D Gaussian avatars into lightweight linear blendshapes, reducing CPU animation costs by up to 1000x and enabling 60fps real-time performance on mobile devices.
ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers
ScholarCatalyst:評估 AI 是否擁有「科學家直覺」的學術文獻檢索基準
ScholarCatalyst is a novel benchmark featuring annotations from 184 lead authors to evaluate whether AI can retrieve key inspiring papers from past literature based only on an initial research question.