Training on data collected from a different policy or source, rather than from the current model's live interactions.