Most memory benchmarks only measure accuracy, but DolphinBench shows that practical agent memory systems must balance three things: how well they retrieve information, how much they cost, and how fast they respond.
DolphinBench is a benchmark that evaluates how well AI agents use long-term memory to complete real-world tasks, not just answer questions. It includes 500k tokens of realistic user messages per persona and measures accuracy, cost, and latency together—revealing tradeoffs that single-metric benchmarks miss.