The process of training a model to predict or infer reward functions from human feedback or demonstrations.