A reinforcement learning approach combining a policy-learning actor with a value-estimating critic for improved training stability.