·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining6h◆Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs6h◆DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions6h◆PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs6h◆Incomplete Prompt Jailbreaks in Large Language Models6h◆Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering6h◆Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants6h◆DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making6h◆Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation6h◆Optimizing Hypergraph-Based RAG: Toward Better Fact Extraction and Chunk Retrieval6h◆DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers6h◆StrideDiffusion: Accelerating Diffusion Models for Time-series Generation6h◆Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs6h◆Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks6h◆Auditing Evidence Use in Medical LLM Diagnosis6h◆Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions6h◆Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority6h◆SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration6h◆From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data6h◆V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure6h◆OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining6h◆Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs6h◆DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions6h◆PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs6h◆Incomplete Prompt Jailbreaks in Large Language Models6h◆Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering6h◆Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants6h◆DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making6h◆Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation6h◆Optimizing Hypergraph-Based RAG: Toward Better Fact Extraction and Chunk Retrieval6h◆DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers6h◆StrideDiffusion: Accelerating Diffusion Models for Time-series Generation6h◆Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs6h◆Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks6h◆Auditing Evidence Use in Medical LLM Diagnosis6h◆Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions6h◆Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority6h◆SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration6h◆From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data6h◆V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure6h◆
News/Auditing Evidence Use in Medical LLM Diagnosis
arxiv
PublishedJuly 24, 2026 at 4:00 AM

Auditing Evidence Use in Medical LLM Diagnosis

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2607.20848v1 Announce Type: new Abstract: Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone does not show whether the model used the case evidence appropriately. We present a behavioral audit of evidence use in medical diagnosis. For e

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivOPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining6harxivStochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs6harxivDecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions6harxivPlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs6h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews