Long-context evaluation needs to measure both accuracy and computational efficiency—the same model can be dramatically more or less efficient depending on its processing strategy, and current benchmarks don't capture this variation.
This paper introduces LongHarness Bench, a benchmark for evaluating how well language models handle long documents using different processing strategies. The benchmark tests both accuracy and efficiency, requiring models to find relevant information across scattered context and reason strategically rather than reading everything.