Comparing multiple responses sampled from the same model to construct rewards from their relative rankings.