·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines11m◆Visual Cue Guided Video Planning for Generalizable Robot Navigation11m◆A unified framework for global and local interpretability using adaptive derivative-ordered random explanation11m◆After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem11m◆Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs11m◆Enhancing Extubation Failure Prediction with LLM-Derived Features from Respiratory Therapy Clinical Notes11m◆Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits11m◆DANTINOX: A Unified Framework for Multi-Paradigm Language Modeling11m◆Think Before You Comfort: Reflective Cognitive Alignment for Protocol-Grounded Elderly Stimulation Agents11m◆Relation Before Entity: Deferred Commitment in Language Model Factual Recall11m◆From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings11m◆MudawanSn: A Gold-Standard Wolof-Arabic Parallel Corpus for Machine Translation11m◆Register Bias in Complexity-Based Large Language Model Routing11m◆Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation11m◆Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant11m◆How AI Assistants Respond to Repeated Abuse11m◆Myovox: Reading Speech from the Muscles of the Face11m◆Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues11m◆No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback11m◆Does Moral Reasoning Training Help or Hurt? Red-Teaming RL-Trained Ethical Agents with Persona Attacks11m◆Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines11m◆Visual Cue Guided Video Planning for Generalizable Robot Navigation11m◆A unified framework for global and local interpretability using adaptive derivative-ordered random explanation11m◆After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem11m◆Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs11m◆Enhancing Extubation Failure Prediction with LLM-Derived Features from Respiratory Therapy Clinical Notes11m◆Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits11m◆DANTINOX: A Unified Framework for Multi-Paradigm Language Modeling11m◆Think Before You Comfort: Reflective Cognitive Alignment for Protocol-Grounded Elderly Stimulation Agents11m◆Relation Before Entity: Deferred Commitment in Language Model Factual Recall11m◆From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings11m◆MudawanSn: A Gold-Standard Wolof-Arabic Parallel Corpus for Machine Translation11m◆Register Bias in Complexity-Based Large Language Model Routing11m◆Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation11m◆Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant11m◆How AI Assistants Respond to Repeated Abuse11m◆Myovox: Reading Speech from the Muscles of the Face11m◆Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues11m◆No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback11m◆Does Moral Reasoning Training Help or Hurt? Red-Teaming RL-Trained Ethical Agents with Persona Attacks11m◆
News/Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning
arxiv
PublishedSeptember 16, 2026 at 4:00 AM
—neutral

Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2609.17386v1 Announce Type: new Abstract: Test-time prompt tuning (TPT) enables adaptation on a single test instance, achieving improved accuracy but often sacrificing calibration performance. Most existing calibration methods introduce additional regularization terms to promote dispersion acr

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivGeospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines11marxivVisual Cue Guided Video Planning for Generalizable Robot Navigation11marxivA unified framework for global and local interpretability using adaptive derivative-ordered random explanation11marxivAfter the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem11m
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews