Safety metrics like toxicity scores can mask real harms—discrimination doesn't disappear during model training, it just becomes harder to detect. Developers need better evaluation methods that catch representational bias, not just explicit toxicity.
This paper reveals that safety improvements in GPT models don't actually reduce gender discrimination—they transform it into subtler forms.