The goal isn't to make models resistant to all external signals, but to teach them when to trust context and when to ignore it—a nuanced skill that requires balanced training across clean, misleading, and irrelevant contexts.
Language models often blindly follow external context, even when it's wrong. This paper introduces MIST, a benchmark that tests when models should trust context versus ignore it, and SCOPE, a training method that teaches models to selectively trust helpful context while rejecting misleading signals—without becoming useless when context is actually correct.