·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
SoftBank’s CEO isn’t the only one with questions about Elon Musk’s orbital data center hype14h◆Margaret Atwood says the problem with AI is ‘garbage in, garbage out’16h◆Apple Vision Pro exec is reportedly leaving for OpenAI18h◆The fittest founder in the room got cancer. Here’s how he used AI to fight back.21h◆Why is Apple asking me to pay more for Big Tech’s AI obsession?22h◆Asian AI startups launch Mythos-like models as Anthropic’s export ban drags on23h◆LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation1d◆Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training1d◆Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems1d◆TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation1d◆An Empirical Study of LLM-Generated Specifications for VeriFast1d◆State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading1d◆Delegation and Verification Under AI1d◆A Concept of Possibility for Real-World Events1d◆Parametric Open Source Games1d◆Semantic Early-Stopping for Iterative LLM Agent Loops1d◆Life After Benchmark Saturation: A Case Study of CORE-Bench1d◆Diagnosing Task Insensitivity in Language Agents1d◆GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning1d◆An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models1d◆SoftBank’s CEO isn’t the only one with questions about Elon Musk’s orbital data center hype14h◆Margaret Atwood says the problem with AI is ‘garbage in, garbage out’16h◆Apple Vision Pro exec is reportedly leaving for OpenAI18h◆The fittest founder in the room got cancer. Here’s how he used AI to fight back.21h◆Why is Apple asking me to pay more for Big Tech’s AI obsession?22h◆Asian AI startups launch Mythos-like models as Anthropic’s export ban drags on23h◆LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation1d◆Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training1d◆Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems1d◆TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation1d◆An Empirical Study of LLM-Generated Specifications for VeriFast1d◆State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading1d◆Delegation and Verification Under AI1d◆A Concept of Possibility for Real-World Events1d◆Parametric Open Source Games1d◆Semantic Early-Stopping for Iterative LLM Agent Loops1d◆Life After Benchmark Saturation: A Case Study of CORE-Bench1d◆Diagnosing Task Insensitivity in Language Agents1d◆GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning1d◆An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models1d◆
News/A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning
arxiv
PublishedApril 28, 2026 at 4:00 AM

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2604.23114v1 Announce Type: new Abstract: In limited-data settings, a single endpoint mean of an evaluation metric such as the Continuous Ranked Probability Score (CRPS) is itself a random variable, yet it is routinely reported as if it were a stable property of the method. We study when this

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivLCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation1darxivHelpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training1darxivInstruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems1darxivTAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation1d
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews