Sparse autoencoders don't compose semantically the way we might hope: their active feature sets track internal model structure rather than human conceptual categories, limiting their usefulness for understanding what language models actually represent.
This paper examines whether sparse autoencoders (SAEs) in language models can capture human-like semantic categories by analyzing which SAE features activate together.