LLMs can learn to teach better by optimizing directly for student learning outcomes rather than following predefined teaching rules—this adaptive approach works across different learner types and aligns with how human teachers actually work.
Sherpa trains LLMs to teach adaptively by using reinforcement learning with multiple simulated student archetypes. Instead of relying on fixed teaching demonstrations, the system directly optimizes for student learning outcomes, enabling teachers to personalize instruction. Results show 20.5 percentage point improvements in student performance and 79.6% human preference over baseline models.