A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control — ThinkLLM