GUI agents can now self-improve after deployment by learning from their own failures through AI-guided reflection and self-distillation, achieving 7.4% accuracy gains without requiring human annotations.
This paper introduces a test-time adaptation framework for GUI visual grounding that allows models to improve after deployment without human feedback. The system uses a closed-loop process: agents explore interfaces, an AI reflector evaluates results and explains failures, and a self-distillation method internalizes these insights back into the model weights.