Most agents show some improvement from retained experience, but the gains are uneven and often don't follow the intended learning pathway—suggesting that simply storing information doesn't guarantee agents will use it effectively to improve.
PAST-Bench is a benchmark that tests whether personal AI agents actually improve over time by learning from retained experience across sessions. The benchmark isolates this capability by running agents through task sequences with experience turned on and off, measuring both performance gains and whether improvements follow the intended save-retrieve-update pathway.