Choosing a representative sample of tasks to evaluate in detail, then using results to estimate performance on the full benchmark.