Small model failures aren't noise—they're structured patterns you can reuse. By studying what weaker models get wrong, you can guide stronger models to avoid similar mistakes with minimal computational overhead.
CritICL uses failure patterns from smaller models to guide larger models during inference. Instead of discarding weak model mistakes, the method captures these predictable failure modes and feeds them as critique examples to help stronger models reason better—achieving better performance than standard methods while using fewer tokens.