arXivAI Research
FERPO: Forward Entropy-Regularized Policy Optimization Without Action Gradients
FERPO:免除動作梯度的前向熵正規化策略最佳化,實現高效且穩定的強化學習
FERPO is a novel online reinforcement learning algorithm that bypasses action-gradient calculations of the critic, using a forward-KL objective to enhance sample efficiency and speed in continuous control.
2 min read