The Loss Does Not See the Basis, but Adam Does

By Devender Singh · Paper · cs.LG

Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under $(U, V) \mapsto (UQ, VQ)$. Gradient

Cs.lg

View original

HomeResourceLoading…