A language model trained to evaluate the quality of another model's outputs and identify areas for improvement.