Training a policy to choose actions based on context, treating each decision point independently rather than optimizing long-term sequences.