arxiv
PublishedMay 27, 2026 at 4:00 AM
MuCon: Clipped Muon Updates for LLM Training
Publisher summary· verbatim
arXiv:2605.26459v1 Announce Type: new Abstract: Muon-style optimizers take a matrix-valued momentum or preconditioned update $B = U \operatorname{diag}(\sigma_1,\ldots,\sigma_r) V^\top$ and replace it with its canonical partial polar factor $\operatorname{Pol}(B) = U V^\top$. This maps every nonzero
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivFrom Words to Widgets for Controllable LLM Generation4harxivScalable Optimal Transport Algorithm for Network Alignment4harxivWhat Does Goodness Measure? A Likelihood-Ratio Account of Forward-Forward Learning4harxivWhen Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary4hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗