ComPO offers a gradient-free alternative to direct preference optimization that may better handle preference pairs with small margins, with both theoretical convergence guarantees and empirical improvements across major LLM families.
This paper introduces ComPO, a new method for aligning language models with human preferences that uses comparison oracles instead of directly optimizing preference losses. Unlike standard approaches, ComPO extracts directional signals from preference pairs without computing gradients, and includes both offline and online variants with theoretical guarantees.