Single stereotype sentence pairs are unreliable for measuring LLM bias; use dual minimal pairs and mutual information-based metrics instead for consistent, language-agnostic bias evaluation.
This paper identifies a flaw in how bias is typically measured in language models: comparing just two sentences about stereotypes can give contradictory results depending on how you rephrase them. The authors propose a better approach using dual comparisons and introduce metrics based on mutual information to more reliably measure stereotype bias across languages and models.