10 / Machine learning Experimental study
First, fix the ruler.
Grokking accelerators against a tuned, sustained baseline.
An apparent speed-up can come from the comparison itself. On the modular-addition testbed, this study retests interventions against a tuned baseline and measures sustained generalisation.
The protocol uses multiple seeds, predictions recorded before computation, and a second independently written testbed. Crossing an accuracy threshold once is distinguished from staying above it.
Where the claim stops.
The conclusions concern the tested grokking setup and interventions. They do not establish that those methods fail on every task. The research repository documents the experiments; a related preprint is also available.
Source: grokking-honest-ruler / README.md. Summary prepared from the local research record, 11 September 2026.
DOI 10.5281/zenodo.22196294 ↗Research led by Aleksei Kudriashov (Alex Komang). Mathematical paper and witness text/data: CC BY 4.0 where stated in the source repository.