Language model agents can adapt to new tasks by learning to revise their execution harness (program structure) rather than their weights, enabling test-time adaptation that generalizes to unseen tasks.
This paper introduces harness learning, a method where an AI agent learns to improve its own executable program (harness) that controls how a language model makes decisions and uses tools. Instead of changing the model's weights, a separate proposer model learns to revise the harness structure based on task feedback.