09 / Machine learning Experimental study
A rule learned. Then erased.
Auxiliary memory, dying gradients and the loss of generalisation.
A learnable per-example memory competes with a grokking network. Once the memory absorbs the training task, the remaining optimisation dynamics can erase a rule the network had already found.
The study separates mechanisms with causal interventions and independently implemented testbeds. A language-model extension finds an important boundary: continued training itself can erase unreinforced knowledge.
Where the claim stops.
A naive protective weight-decay schedule did not reliably preserve the rule. The mechanism and that negative result are both retained. The research repository documents the experiments; a related preprint is also available.
Source: grokking-race-eraser / README.md. Summary prepared from the local research record, 11 September 2026.
DOI 10.5281/zenodo.22196294 ↗Research led by Aleksei Kudriashov (Alex Komang). Mathematical paper and witness text/data: CC BY 4.0 where stated in the source repository.