Sliding-window constraints make online learning fundamentally harder than offline optimization, but rare policy updates combined with optimistic planning can achieve sublinear regret while maintaining exact feasibility.
This paper studies how to make optimal decisions in linear bandit problems when actions must satisfy strict sliding-window constraints—meaning every consecutive block of actions must come from an allowed set.