User feedback is a strong improvement signal for LLMs, but standard LLM-based evaluation systematically fails to recognize when models successfully apply it—meaning we're underestimating feedback's real value.
This paper shows that user feedback is actually a valuable signal for improving LLMs, but current evaluation methods fail to recognize when models successfully use it.