Training technique where intermediate layer outputs are supervised to predict the next token, enabling models to work at multiple depths.