A model can detect when its own context has become unreliable by monitoring prediction error—use this self-awareness to filter bad context before making decisions, rather than blindly trusting a critic's action choices.
Decision Transformers struggle on long rollouts because their conditioning context drifts from training data. This paper proposes Trust Guided Decision Transformer (TGDT), which uses the model's own prediction errors to identify when context becomes unreliable, then filters out bad context before selecting actions.