Self-repair in language models isn't adaptive compensation—it's pre-existing counterweights responding predictably to ablation. You can predict how a component will respond to intervention from its fixed weights alone.
When you disable a component in a language model, other parts often seem to compensate—a phenomenon called 'self-repair.' This paper shows it's not actually repair: it's pre-existing counterweights doing their normal job.