A constraint that limits how much a policy can change during training to prevent excessive drift from the original model.