·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs7h◆Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs7h◆A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies7h◆SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification7h◆Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs7h◆ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents7h◆Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems7h◆Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers7h◆Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains7h◆The Curse of Multilinguality in Lexical Normalization7h◆RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces7h◆The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space7h◆TempCloze: Can Video-LLMs Identify the Missing Middle?7h◆Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models7h◆SARTM: Segment Any RGB Thermal Model with Language aided Distillation7h◆Unsupervised Partner Design Enables Robust Ad-hoc Teamwork7h◆Agentic Empirical Asset Pricing: Methodological Foundations7h◆Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement7h◆BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection7h◆FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study7h◆Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs7h◆Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs7h◆A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies7h◆SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification7h◆Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs7h◆ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents7h◆Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems7h◆Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers7h◆Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains7h◆The Curse of Multilinguality in Lexical Normalization7h◆RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces7h◆The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space7h◆TempCloze: Can Video-LLMs Identify the Missing Middle?7h◆Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models7h◆SARTM: Segment Any RGB Thermal Model with Language aided Distillation7h◆Unsupervised Partner Design Enables Robust Ad-hoc Teamwork7h◆Agentic Empirical Asset Pricing: Methodological Foundations7h◆Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement7h◆BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection7h◆FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study7h◆
News/ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection
arxiv
PublishedMay 14, 2026 at 4:00 AM
▲bullish

ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2605.12879v1 Announce Type: new Abstract: Doubly-stochastic attention has emerged as a transport-based alternative to row-softmax attention, with recent Transformer variants using it to reduce attention sinks and rank collapse while improving performance. In this family, the standard approach

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
02
  • 01
    Sinkhorn
  • 02
    ASAP
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#transformer#attention#machine-learning#optimization

No replies yet. Be first.

Mentioned models
02
  • 01
    Sinkhorn
  • 02
    ASAP
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#transformer#attention#machine-learning#optimization

Related coverage

More from ARXIV
arxivResidual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs7harxivControl-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs7harxivA Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies7harxivSOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification7h
The Bubble Brief
WEEKLY

Read transformer insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews