Wasserstein policy gradient converges exponentially fast for entropy-regularized linear-quadratic control without exponential slowdown as regularization decreases, making it a theoretically sound alternative to standard policy gradient methods.
This paper studies how to optimize control policies using Wasserstein gradients—a method that updates action distributions by moving them in action space. For linear-quadratic control problems with entropy regularization, the authors prove that this approach reduces to a simple finite-dimensional system that converges reliably to the optimal policy, even as the regularization strength changes.