·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Perplexity trusts GPT-6 Astra with end-to-end systems-2485m◆TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents2h◆Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision2h◆Monadic Second-Order Logic in HOL: Deep and Shallow with Automated Faithfulness (Extended Preprint)2h◆Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation2h◆ScaleResfusion: Residual Rectified Flow based on Residual Vector Field2h◆SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models2h◆Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy2h◆Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements2h◆Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning2h◆ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks2h◆PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving2h◆Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement2h◆Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech2h◆KuaiRP Series Role-playing Models Technical Report2h◆A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies2h◆RetroThinker: Enabling Retrospective Thinking in Speech LLMs2h◆The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement2h◆Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens2h◆CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models2h◆Perplexity trusts GPT-6 Astra with end-to-end systems-2485m◆TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents2h◆Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision2h◆Monadic Second-Order Logic in HOL: Deep and Shallow with Automated Faithfulness (Extended Preprint)2h◆Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation2h◆ScaleResfusion: Residual Rectified Flow based on Residual Vector Field2h◆SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models2h◆Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy2h◆Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements2h◆Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning2h◆ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks2h◆PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving2h◆Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement2h◆Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech2h◆KuaiRP Series Role-playing Models Technical Report2h◆A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies2h◆RetroThinker: Enabling Retrospective Thinking in Speech LLMs2h◆The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement2h◆Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens2h◆CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models2h◆
News/Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit
arxiv
PublishedApril 21, 2026 at 4:00 AM
▲bullish

Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2604.15356v1 Announce Type: cross Abstract: Recent work on KV cache quantization, culminating in TurboQuant, has approached the Shannon entropy limit for per-vector compression of transformer key-value caches. We observe that this limit applies to a strictly weaker problem than the one that ac

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
01
  • 01
    TurboQuant
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#compression#quantization#transformers#information-theory

No replies yet. Be first.

Mentioned models
01
  • 01
    TurboQuant
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#compression#quantization#transformers#information-theory

Related coverage

More from ARXIV
arxivTeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents2harxivAmbient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision2harxivMonadic Second-Order Logic in HOL: Deep and Shallow with Automated Faithfulness (Extended Preprint)2harxivNatural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation2h
The Bubble Brief
WEEKLY

Read compression insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews