Optimizing search query understanding components individually with search-engine-derived rewards outperforms single end-to-end optimization, showing that understanding how each component affects downstream retrieval matters more than just matching labels.
This paper presents a reinforcement learning framework for query understanding in search systems that optimizes multiple components (like intent classification and query expansion) separately using rewards from live search engine interactions, rather than treating it as a single end-to-end problem.