By separating question sources from evaluation criteria in research papers, you can create reliable reward signals for training AI on research planning—without the model just paraphrasing its way to high scores.
PaperGym turns research papers into training environments for AI scientists by extracting evaluation criteria (rubrics) from paper structure—questions from goals/background, criteria from methods/experiments. This solves the problem that research planning has no verifiable answers, enabling reinforcement learning training.