A preprint posted to arXiv on October 2, 2026, proposes Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm for continuous control.

Continuous-control RL often relies on action gradients of a learned critic. The arXiv CS.LG preprint says critics are trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives. FERPO avoids differentiating the critic with respect to actions by deriving an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler divergence, then fitting the actor to that target via a forward-KL objective and self-normalized importance sampling.

The KL regularization limits the target distribution's deviation from the rollout policy to keep importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of modes, the forward-KL objective encourages coverage of multiple high-value modes and promotes exploration, according to the paper.

The authors report competitive performance and sample-efficiency gains on MuJoCo Playground and ManiSkill, and computational benchmarks show faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO). The preprint has not been peer reviewed.