Artificially generated pairs of responses (good and bad) used to validate whether an evaluation criterion is effective.