08 / Machine learning Experimental study
The structure of what we learn.
Data selection, spectral predictors and conditional generalisation.
At a fixed training-set size, different arrangements of modular-addition examples can lead to different grokking behaviour. The study tests a spectral predictor computed before training.
Experiments on small language models then examine coverage, composition and window freshness. Their effects depend on the corpus: a result on heterogeneous text need not carry over to homogeneous prose.
Where the claim stops.
This is a conditional empirical finding on the tested tasks and models, not a universal data-selection law. The research repository documents the experiments; a related preprint is also available.
Source: grokking-data-law / README.md. Summary prepared from the local research record, 11 September 2026.
DOI 10.5281/zenodo.22196294 ↗Research led by Aleksei Kudriashov (Alex Komang). Mathematical paper and witness text/data: CC BY 4.0 where stated in the source repository.