Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

By Ali Janati · Paper · cs.AI

Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seeds the selected AdamW reference

Grok · Cs.ai

View original

HomeResourceLoading…