You can measure training data importance during pretraining by tracking parameter trajectories, revealing that different data types matter at different training stages—literature early, STEM later—without needing task-specific validation sets.
This paper introduces a new method to measure how much training data influences language model development without needing to pick specific downstream tasks. Instead of testing on particular benchmarks, the researchers measure influence by tracking how each piece of training data pushes the model toward its final parameters.