Using a fixed set of reference responses as anchors to rank new outputs and derive rewards efficiently.