arxiv
PublishedApril 28, 2026 at 4:00 AM
A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning
Publisher summary· verbatim
arXiv:2604.23114v1 Announce Type: new Abstract: In limited-data settings, a single endpoint mean of an evaluation metric such as the Continuous Ranked Probability Score (CRPS) is itself a random variable, yet it is routinely reported as if it were a stable property of the method. We study when this
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivLCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation1darxivHelpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training1darxivInstruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems1darxivTAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation1dThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗