nusadua.dev← All researchIndependent by design

09 / Machine learning Experimental study

A rule learned. Then erased.

Auxiliary memory, dying gradients and the loss of generalisation.

A learnable per-example memory competes with a grokking network. Once the memory absorbs the training task, the remaining optimisation dynamics can erase a rule the network had already found.

The study separates mechanisms with causal interventions and independently implemented testbeds. A language-model extension finds an important boundary: continued training itself can erase unreinforced knowledge.

Where the claim stops.

A naive protective weight-decay schedule did not reliably preserve the rule. The mechanism and that negative result are both retained. The research repository documents the experiments; a related preprint is also available.

Source: grokking-race-eraser / README.md. Summary prepared from the local research record, 11 September 2026.

DOI 10.5281/zenodo.22196294 ↗

Research led by Aleksei Kudriashov (Alex Komang). Mathematical paper and witness text/data: CC BY 4.0 where stated in the source repository.

Keep following the questionsFirst, fix the ruler. ↗