LLMs maintain readable and writable internal user models that directly influence safety behavior; different models independently converge on similar user representations, suggesting this is a fundamental property of how language models condition their responses.
This paper introduces Belief Self-Distillation (BSD), a technique to extract and manipulate how LLMs represent their users internally.