Gradient information computed only at a model's final layer, cheaper to compute than full-parameter gradients.