Assessing AI agent performance by having original researchers grade the agent's work on their own unpublished research.