·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models4h◆Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders4h◆Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA4h◆Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging4h◆Multi-Mask Diffusion Language Models for Few-Step Generation4h◆Solar Open 2 Technical Report4h◆The Geometry of Personality: Activation Steering with Jungian Cognitive Functions4h◆Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning4h◆H$^2$SD: Hybrid Hindsight Self-Distillation4h◆LunarFM: A Shared Multimodal Representation of the Moon's Surface4h◆Prior laundering: learned priors with inherited, undetectable overconfidence4h◆Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery4h◆Prompt as a Data Type: In-Database LLM Prompt Management and Rewriting4h◆CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference4h◆Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support4h◆Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models4h◆A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data4h◆Simpson's Paradox in Behavioral Curves: How Aggregation Distorts Parametric Models of User Dynamics4h◆Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO4h◆Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning4h◆A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models4h◆Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders4h◆Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA4h◆Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging4h◆Multi-Mask Diffusion Language Models for Few-Step Generation4h◆Solar Open 2 Technical Report4h◆The Geometry of Personality: Activation Steering with Jungian Cognitive Functions4h◆Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning4h◆H$^2$SD: Hybrid Hindsight Self-Distillation4h◆LunarFM: A Shared Multimodal Representation of the Moon's Surface4h◆Prior laundering: learned priors with inherited, undetectable overconfidence4h◆Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery4h◆Prompt as a Data Type: In-Database LLM Prompt Management and Rewriting4h◆CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference4h◆Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support4h◆Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models4h◆A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data4h◆Simpson's Paradox in Behavioral Curves: How Aggregation Distorts Parametric Models of User Dynamics4h◆Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO4h◆Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning4h◆
News/SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
arxiv
PublishedApril 10, 2026 at 4:00 AM
▲bullish

SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2604.07663v1 Announce Type: new Abstract: The AdamW optimizer, while standard for LLM pretraining, is a critical memory bottleneck, consuming optimizer states equivalent to twice the model's size. Although light-state optimizers like SinkGD attempt to address this issue, we identify the embedd

Models mentioned
01
  • 01meta-llama logo
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
01
  • 01
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
03
#optimization#memory-efficiency#large-language-models
Mentioned companies
01
Meta

No replies yet. Be first.

Mentioned models
01
  • 01
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
03
#optimization#memory-efficiency#large-language-models
Mentioned companies
01
Meta

Related coverage

More from ARXIV
arxivA Consensus-Based Framework for Relative Preference Evaluation of Large Language Models4harxivProbing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders4harxivData Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA4harxivEnjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging4h
The Bubble Brief
WEEKLY

Read optimization insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews