LLM judges used to score AI outputs are not stable measurement instruments: identical requests produce different rankings across days and providers, breaking the assumption that 'same model name = same scorer.
This paper audits whether language models used as judges produce consistent scores across repeated requests—a critical assumption for using them to evaluate AI systems. Testing over 52,000 requests, the authors found that identical inputs returned different rankings on different days (agreement of 0.78 when 0.99 was required), making LLM judges unreliable measurement instruments.