LLMs don't just detect moral content; they organize it geometrically with a shared moral component and independent dimensions per foundation, meaning they capture nuanced relationships between different ethical frameworks and represent unresolved moral tensions.
This paper investigates how large language models organize moral knowledge by training linear probes on six moral foundation categories and analyzing their geometric relationships in representation space.