Training on a carefully curated, grade-level-appropriate curriculum creates a sandbox for studying knowledge acquisition with clear boundaries—useful for understanding how models learn and what happens when you try to teach them new concepts.
Researchers created LittleLeaner, a 5B-parameter language model trained on an 88B-token curriculum limited to U.S. Grade 5 material, to study how models acquire knowledge under controlled conditions. Unlike models trained on messy web data, LittleLeaner has clear, interpretable knowledge boundaries, making it easier to understand what the model knows and how it learns new information.