Using forward-KL instead of reverse-KL for policy fitting encourages broader exploration of high-value actions and avoids the computational cost of differentiating critics, leading to faster and more sample-efficient learning.
FERPO is a reinforcement learning algorithm that improves policies by using critic values directly rather than differentiating through them. It derives optimal target actions using entropy regularization and fits the actor to these targets using forward-KL divergence, which encourages exploring multiple high-value action modes while keeping importance weights stable.