Language models struggle with WiC not because they can't understand word meanings, but because they lack clear guidance on what level of semantic detail matters—adding explicit sense options fixes this and reveals models often overthink distinctions.
This paper shows that Word-in-Context (WiC) tasks are harder for language models than traditional Word Sense Disambiguation (WSD) because WiC lacks an explicit sense inventory. By providing candidate senses to models, performance improves significantly, and human evaluation reveals many errors stem from disagreement about sense granularity rather than true comprehension failures.