·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal4h◆RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching4h◆VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs4h◆Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models4h◆An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism4h◆Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching4h◆Candidate Attended Dialogue State Tracking Using BERT4h◆An Exam for Active Observers4h◆A Neuro-Symbolic Approach for Probabilistic Reasoning on Graph Data4h◆Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces4h◆Evidence-Aware MapReduce for Forkable Compute4h◆From Plausible to Actionable: A Position on LLM Self-Explanations4h◆Robust Explanations for User Trust in Enterprise NLP Systems4h◆Brain-CLIPLM: Semantic Compression for EEG-to-Text Decoding4h◆Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery4h◆Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers4h◆SODA: Semi On-Policy Black-Box Distillation for Large Language Models4h◆Graph Coloring Approach to Solving Sudoku with Oscillatory Neural Networks4h◆MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation4h◆Retraining Seeks Stable Signals4h◆Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal4h◆RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching4h◆VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs4h◆Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models4h◆An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism4h◆Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching4h◆Candidate Attended Dialogue State Tracking Using BERT4h◆An Exam for Active Observers4h◆A Neuro-Symbolic Approach for Probabilistic Reasoning on Graph Data4h◆Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces4h◆Evidence-Aware MapReduce for Forkable Compute4h◆From Plausible to Actionable: A Position on LLM Self-Explanations4h◆Robust Explanations for User Trust in Enterprise NLP Systems4h◆Brain-CLIPLM: Semantic Compression for EEG-to-Text Decoding4h◆Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery4h◆Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers4h◆SODA: Semi On-Policy Black-Box Distillation for Large Language Models4h◆Graph Coloring Approach to Solving Sudoku with Oscillatory Neural Networks4h◆MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation4h◆Retraining Seeks Stable Signals4h◆
News/From SGD to Muon: Adaptive Optimization via Schatten-p Norms
arxiv
PublishedMay 21, 2026 at 4:00 AM
▲bullish

From SGD to Muon: Adaptive Optimization via Schatten-p Norms

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2605.19781v1 Announce Type: new Abstract: Modern optimizers, like Muon, impose matrix-wise geometry constraints on their updates. These matrix-wise constraints can be unified under Linear Minimization Oracle (LMO) theory. However, all current methods impose fixed LMO geometries for the update

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
05
  • 01
    Muon
  • 02
    SGD
  • 03
    Adam
  • 04
    AdamW
  • 05
    MuAdam
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#optimization#deep learning#neural networks#artificial intelligence

No replies yet. Be first.

Mentioned models
05
  • 01
    Muon
  • 02
    SGD
  • 03
    Adam
  • 04
    AdamW
  • 05
    MuAdam
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#optimization#deep learning#neural networks#artificial intelligence

Related coverage

More from ARXIV
arxivBeyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal4harxivRobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching4harxivVarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs4harxivAdaptive Multi-Step Lookahead Decoding for Diffusion Language Models4h
The Bubble Brief
WEEKLY

Read optimization insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews