Model activation patterns can reliably predict how training on one value transfers to untrained values—a capability that could guide more robust multi-value alignment training without extensive testing.
This paper studies how fine-tuning LLMs on specific values (like honesty or helpfulness) affects their behavior on other unseen values. The authors develop methods to predict these generalization effects using model activation patterns, finding that activation-based representations vastly outperform text-based approaches.