Unified models can learn to self-correct their outputs by applying RL to complete reflection loops, where both the reasoning about what's wrong and the actual image fixes improve together without needing external verifiers.
This paper presents UMM-Reflection, a method that teaches unified multimodal models to critique and fix their own image generations through reinforcement learning. Instead of just generating images once, the model can now look at what it created, identify problems, revise the image, and repeat—all within a single model.