You can build effective evaluation rubrics automatically from scratch by testing whether proposed criteria actually discriminate between high and low-quality responses, without needing human-written examples or preference data.
This paper introduces a method to automatically generate evaluation rubrics (scoring criteria) for LLMs using only a query and no human annotations. It works by proposing candidate rubrics, testing them against synthetic response pairs to verify they actually distinguish good answers from bad ones, and keeping only the useful ones.