Instruction tuning makes models sound more confident without improving accuracy, and it reduces the diversity of explanations they provide—a potential concern for transparency and reliability.
This paper investigates how instruction tuning affects language models' confidence levels and the diversity of their explanations. The researchers found that instruction tuning makes models express higher confidence in their answers, but this doesn't match improvements in actual accuracy.