You can improve a frozen language model's performance on specialized tasks by treating policy refinement as a human-in-the-loop process: have an AI critic identify recurring failures, propose natural-language policy changes, and let domain experts decide what gets deployed.
This paper presents Policy Iteration with Human Feedback (PIHF), a method that improves a fixed language model's performance on rare-disease diagnosis by iteratively refining its decision-making policy through human expert review.