When deploying LLMs for decisions without ground truth, you can assess individual recommendations by measuring preference stability and scope—not just overall model reliability—using a practical four-tier certification system.
This paper introduces 'epistemic warrant,' a framework for assessing whether to trust individual LLM recommendations when you can't verify the answer. Instead of evaluating broad model properties, the authors create a four-tier system that measures how stable a model's preference is and how widely it applies, validated through expert and crowd-worker agreement.