The degree to which a benchmark tests different strategies or behaviors an agent might use to solve tasks.