When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning — ThinkLLM