AI safety evaluations that only report accuracy on a single API call miss critical behavioral variations—modality, search conditions, and response consistency matter significantly and can even reverse performance rankings between deployment methods.
This paper reveals critical gaps in how AI safety benchmarks are evaluated by testing ChatGPT across different access methods (chat UI vs API), with and without web search, across multiple runs.