You can prevent large models from forgetting old skills during new training by using a learnable gate that only activates weight updates when the input matches the current training distribution—no need to store old data.
This paper addresses catastrophic forgetting in large language models by treating it as a geometric problem in weight space.