Don't trust tool-use benchmarks for small models—use verbatim-reproduction checks and token-probability probes to verify genuine capability before claiming tool use works.
Small language models can appear to use tools correctly on standard benchmarks while actually just memorizing training examples. This paper shows how keyword-matching tests miss real failures, proposes cheap diagnostic checks to catch false positives, and demonstrates a targeted fix that repairs tool-use ability in a 1.1B parameter model using minimal compute.