Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
By Zihan Liu · Paper · cs.LG
We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse throughout training despite substa