Attention heads need both the ability to abstain from attending and to filter noise from values—their importance shifts with model scale, suggesting future architectures should support both primitives.
This paper identifies two missing capabilities in standard softmax attention: abstention (allowing heads to output nothing instead of always producing weighted combinations) and noise filtering (suppressing interference from mixed features).