Embedding models need explicit training with distractors to reliably follow retrieval instructions—current models are brittle to irrelevant query information despite appearing to understand instructions.
This paper investigates why embedding models struggle to follow retrieval instructions, even simple ones. The researchers discovered that models fail when distractor information is present in queries and show that fine-tuning with query-side distractors significantly improves instruction-following ability without hurting performance on other tasks.