2 papers
cs.AI2026
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau +1
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on…
cs.LG2026
Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers
Andrei Kanavalau, Carmen Amo Alonso, Sanjay Lall
Normalization layers are standard in transformers, but it is not clear whether their sample-dependent computations are necessary throughout both training and inference. This work d…