A reinforcement learning method that updates value estimates using the difference between predicted and observed rewards.