arxiv
PublishedSeptember 7, 2026 at 4:00 AM
—neutral
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Publisher summary· verbatim
arXiv:2609.05275v1 Announce Type: new Abstract: Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly l
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivExtremely Sparse Supervision Incentivizes Reasoning Ability22harxivConstructing and Evaluating Clinical Reasoning Trajectories for Medical Agent22harxivWhen Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs22harxivBlockchain-Enabled Secure Logging for Fiscal Electronic Mechanisms: Evaluation of the Greek eSEND and myDATA Tax Systems22hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗