When evaluating whether models improve on individual problems, you need a separately measured null baseline for every statistic—not just comparing two noisy estimates—or you'll mistake measurement artifacts for real capability gains.
This paper audits claims about language model self-improvement by comparing a fine-tuned model against a frozen control run through identical evaluation pipelines. The authors identify seven measurement artifacts that flip reported findings, showing that many apparent capability gains are statistical illusions from batching effects and noisy comparisons rather than real improvements.