A fine-tuning technique that aligns a model's behavior with human preferences by directly optimizing based on preference data rather than using reinforcement learning.
Adhering to complex, structured, or constrained instructions