CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes — ThinkLLM