When Does Muon Help Agentic Reinforcement Learning?

By Kai Ruan · Paper · cs.LG

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. U

AI Agents · Cs.lg

View original

HomeResourceLoading…