Techniques like batch norm or layer norm that rescale activations to have consistent statistics across training.