Decision models are surprisingly brittle: short, natural-sounding context additions can redirect correct predictions to high-confidence wrong answers, suggesting their probability outputs shouldn't be trusted as reliable decision interfaces without additional safeguards.
This paper reveals that decision models like Jev—which map text to probability distributions over choices—are vulnerable to subtle context manipulation. Researchers show that naturally-written contextual additions can flip correct decisions to wrong ones in 61% of cases, even when the original question and answer remain unchanged.