arxiv
PublishedJune 10, 2026 at 4:00 AM
AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping
Publisher summary· verbatim
arXiv:2502.11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause of loss spikes by investigating individual factors, we observe that, in practice, such spikes are t
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivCost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems13harxivSelf-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems13harxivLearning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance13harxivThe Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests13hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗