Approximate Muon with low-rank adapters

By Ben Anson · Paper · cs.LG

The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common PEFT method, LoRA, does not naturally combine with Muon since

Cs.lg

View original

HomeResourceLoading…