When evaluating ML for critical infrastructure like power grids, the evaluation setup matters as much as the model—standardizing how you measure performance is essential for meaningful comparisons and real-world deployment.
This paper proposes a standardized framework for evaluating machine learning models in power system protection, addressing the problem that near-perfect reported scores often depend on unstated evaluation choices.