arxiv
PublishedJune 6, 2026 at 4:00 AM
—neutral
Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction
Publisher summary· verbatim
arXiv:2606.05863v1 Announce Type: cross Abstract: Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales. We formalize this phenomenon by separating the fast decay of the classification loss from the slower simplification of the lear
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivThe Steering Budget: Examples beat Knobs18harxivPolestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs18harxivRxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination18harxivWhen a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models18hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗