Training world models to discriminate between alternative action outcomes (rather than just predict accurately) makes them better at helping agents choose the right action when browsing the web.
This paper improves web agents by training world models that predict web states in a way that helps rank candidate actions. Instead of predicting states accurately for their own sake, the model learns to make predicted states distinguishable from each other—so a ranker can tell which action leads to the best outcome.