A learned component that decides, for each token, how many processing passes it should receive to optimize accuracy-compute tradeoffs.