A standard test used to measure how well an AI model performs, which can embed biases about what counts as good output.