·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
EDRAC: Benchmarking Arabic Dialect Reading Comprehension3h◆StainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization3h◆Persistent Entropy as a Detector of Phase Transitions3h◆Why Fine-Tuning Encourages Hallucinations and How to Fix It3h◆Skill Reuse as Compression in Agentic RL3h◆Enabling KV Caching of Shared Prefix for Diffusion Language Models3h◆Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts3h◆Two locked tests of phase-structure features for transition prediction3h◆StudentSim: Training LLM-based Student Simulators3h◆AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes3h◆Aligned but Flattened: Analyzing the Trade-off between Cultural Alignment and Diversity in LLMs3h◆Controllable Image Captioning with Prompt-Conditioned Scene Rewards3h◆RECAP: Regression Evaluation for Continual Adaptation of Prompts3h◆VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows3h◆CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN3h◆HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields3h◆Verdict Instability of OOD Scores under Reference Resampling3h◆The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow3h◆SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations3h◆Position: Privacy Is a Claim, Not a Property of Synthetic Data3h◆EDRAC: Benchmarking Arabic Dialect Reading Comprehension3h◆StainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization3h◆Persistent Entropy as a Detector of Phase Transitions3h◆Why Fine-Tuning Encourages Hallucinations and How to Fix It3h◆Skill Reuse as Compression in Agentic RL3h◆Enabling KV Caching of Shared Prefix for Diffusion Language Models3h◆Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts3h◆Two locked tests of phase-structure features for transition prediction3h◆StudentSim: Training LLM-based Student Simulators3h◆AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes3h◆Aligned but Flattened: Analyzing the Trade-off between Cultural Alignment and Diversity in LLMs3h◆Controllable Image Captioning with Prompt-Conditioned Scene Rewards3h◆RECAP: Regression Evaluation for Continual Adaptation of Prompts3h◆VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows3h◆CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN3h◆HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields3h◆Verdict Instability of OOD Scores under Reference Resampling3h◆The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow3h◆SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations3h◆Position: Privacy Is a Claim, Not a Property of Synthetic Data3h◆
News/Gemma 4, Phi-4, and Qwen3: Accuracy-Efficiency Tradeoffs in Dense and MoE Reasoning Language Models
arxiv
PublishedApril 9, 2026 at 4:00 AM
—neutral

Gemma 4, Phi-4, and Qwen3: Accuracy-Efficiency Tradeoffs in Dense and MoE Reasoning Language Models

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2604.07035v1 Announce Type: new Abstract: Mixture-of-experts (MoE) language models are often expected to offer better quality-efficiency tradeoffs than dense models because only a subset of parameters is activated per token, but the practical value of that advantage depends on end-to-end behav

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivEDRAC: Benchmarking Arabic Dialect Reading Comprehension3harxivStainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization3harxivPersistent Entropy as a Detector of Phase Transitions3harxivWhy Fine-Tuning Encourages Hallucinations and How to Fix It3h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews