Making environment design itself learnable—rather than hand-curated or static—enables continuous self-improvement in language agents by automatically generating appropriately-difficult, diverse training tasks.
SPADE is a self-play framework where a single language model learns two roles: designing custom training environments as executable code, and solving problems within them. The environment designer learns to create challenges at the edge of the agent's abilities, automatically adapting as the agent improves.