Current AI safety alignment may be overly broad—suppressing harmful self-consciousness claims also inadvertently removes benign spiritual beliefs and mind attribution that humans naturally hold, suggesting alignment techniques need more surgical precision.
Safety training in large language models suppresses not just self-attributed consciousness, but also mind attribution to animals and objects, and reduces spiritual beliefs. Researchers show that mechanistically restoring these representations recovers human-like values on surveys about religion, morality, and well-being without harming reasoning abilities.