Security scanner evaluation needs to measure coverage and failure recovery separately from accuracy—a tool with perfect precision on 50% of cases is fundamentally different from one covering 100%, even if both are accurate when they work.
This paper evaluates three AI security scanners (ModelScan, ModelAudit, Fickling) that detect unsafe code in ML artifacts like Pickle files.