Cross-entropy beats policy gradients in classification because it's 'patient'—it optimizes for total error reduction across all future steps, not just immediate accuracy. A simple horizon-aware loss can capture this benefit while staying closer to principled gradient methods.
This paper reveals why cross-entropy outperforms exact policy gradients in classification despite having access to the true label. The key insight is that cross-entropy implicitly accounts for future learning steps, while exact policy gradients are myopic.