A framework for sampling from complex distributions by learning a policy that generates trajectories proportional to a reward signal.