·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval37m◆MasterControl Seventeen Every Time37m◆Speculative Macro Commit for Faster Tool-Using Agents37m◆Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory37m◆A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant37m◆Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation37m◆Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection37m◆DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents37m◆Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents37m◆Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty37m◆AutoGraphForge: Towards Automated Graph Theory Discovery37m◆Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models37m◆GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving37m◆PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing37m◆What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation37m◆CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning37m◆NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis37m◆Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation37m◆Dalek: A Constructive Agent Machine37m◆GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis37m◆ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval37m◆MasterControl Seventeen Every Time37m◆Speculative Macro Commit for Faster Tool-Using Agents37m◆Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory37m◆A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant37m◆Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation37m◆Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection37m◆DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents37m◆Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents37m◆Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty37m◆AutoGraphForge: Towards Automated Graph Theory Discovery37m◆Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models37m◆GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving37m◆PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing37m◆What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation37m◆CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning37m◆NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis37m◆Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation37m◆Dalek: A Constructive Agent Machine37m◆GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis37m◆
News/NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems
arxiv
PublishedJune 12, 2026 at 4:00 AM
—neutral

NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2601.11004v3 Announce Type: replace Abstract: Accurately assessing model confidence is essential for deploying large language models (LLMs) in mission-critical factual domains. While retrieval-augmented generation (RAG) is widely adopted to improve grounding, confidence calibration in RAG sett

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval37marxivMasterControl Seventeen Every Time37marxivSpeculative Macro Commit for Faster Tool-Using Agents37marxivFresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory37m
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews