RL technique that estimates action advantage (Q-value minus baseline) to reduce variance in policy gradient learning.