By diagnosing model weaknesses at the granular level of individual rubric criteria rather than whole prompts, you can generate more targeted fine-tuning data that produces measurably better models across multiple domains.
CRAFT is a diagnostic method that analyzes rubric-based evaluations to identify specific capability gaps in language models, then uses those insights to generate targeted fine-tuning data.