Assessing model outputs using a detailed scoring guide with specific criteria rather than simple right/wrong answers.