Deciding at each generated token whether to continue with a small model or request help from a larger model.