The process of predicting and producing the next word or subword (token) in a sequence, one at a time, during model inference.