Even advanced AI agents fail at retrieving papers that inspired real research (max 0.51 recall), revealing a critical gap in how models search scientific literature—this task requires something beyond current retrieval and reasoning approaches.
ScholarCatalyst is a benchmark dataset where 184 computer science researchers labeled which prior papers inspired their completed projects. The benchmark tests whether AI systems can retrieve these influential papers given only an initial research question and literature available at project start.