Anomaly detection benchmarks hide critical methodological choices; MIRTO shows that reported performance differences between methods often reflect evaluation setup rather than actual capability differences, and that registration alignment and threshold selection are major hidden sources of variance.
MIRTO is an evaluation protocol that makes explicit the hidden choices in unsupervised brain MRI anomaly detection—how maps are aligned, thresholds set, and metrics calculated.