Procedural datasets for training reasoning models need careful design beyond correctness: difficulty calibration, compact targets, and rigorous auditing (combining model review, human judgment, and testing) significantly impact training utility.
This paper introduces Reasoning Core, a collection of 50 procedural generators that create verifiable reasoning problems across diverse domains (math, logic, planning, code, etc.). The authors compare their dataset with three alternatives using completion-supervised fine-tuning on 3B models, showing Reasoning Core achieves better performance on reasoning benchmarks.