arxiv
PublishedApril 15, 2026 at 4:00 AM
—neutral
On the Convergence Analysis of Muon
Publisher summary· verbatim
arXiv:2505.23737v2 Announce Type: replace-cross Abstract: The majority of parameters in neural networks are naturally represented as matrices. However, most commonly used optimizers treat these matrix parameters as flattened vectors during optimization, potentially overlooking their inherent structu
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9harxivSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks9harxivDistribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts9harxivIn RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning9hThe Bubble Brief
WEEKLYRead optimization insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗