11 / Machine learning Experimental study
Give the experiment time.
Activation, final LayerNorm and the training-budget effect.
Short experiments can make an architectural difference look permanent. On the tested character-level models, the ReLU–GELU gap shrinks as training continues, while removing the final LayerNorm retains a cost.
Budget sweeps and two independently written testbeds help distinguish an optimisation-horizon effect from a persistent architectural change.
Where the claim stops.
These are small-model experimental results with stated budgets and seeds. The research repository documents the experiments; a related preprint is also available.
Source: grokking-activation-ln / README.md. Summary prepared from the local research record, 11 September 2026.
DOI 10.5281/zenodo.22196294 ↗Research led by Aleksei Kudriashov (Alex Komang). Mathematical paper and witness text/data: CC BY 4.0 where stated in the source repository.