·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
What to watch for after Jensen Huang’s Japan visit1h◆Can an Apple lawsuit derail OpenAI’s hardware plans?3h◆I hate that I don’t hate this song made with Suno5h◆‘Odyssey’ director Christopher Nolan calls AI an obvious ‘Trojan horse’8h◆Nonprofit Current AI is racing to build the World Wide Web of AI, free for all8h◆Dave Eggers told OpenAI staff that ChatGPT was ‘silencing an entire generation’1d◆Kimi: Threat or menace?1d◆The apps, gadgets, and tools every reader needs1d◆Neil Rimer thinks the AI money is coming back out1d◆The Steering Budget: Examples beat Knobs1d◆Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs1d◆RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination1d◆HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization1d◆When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models1d◆SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation1d◆ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System1d◆LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets1d◆Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility1d◆Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation1d◆LIGO-PINN: Learned Initialization via Gated Optimization to Alleviate Convergence Failures in Physics Informed Neural Networks1d◆What to watch for after Jensen Huang’s Japan visit1h◆Can an Apple lawsuit derail OpenAI’s hardware plans?3h◆I hate that I don’t hate this song made with Suno5h◆‘Odyssey’ director Christopher Nolan calls AI an obvious ‘Trojan horse’8h◆Nonprofit Current AI is racing to build the World Wide Web of AI, free for all8h◆Dave Eggers told OpenAI staff that ChatGPT was ‘silencing an entire generation’1d◆Kimi: Threat or menace?1d◆The apps, gadgets, and tools every reader needs1d◆Neil Rimer thinks the AI money is coming back out1d◆The Steering Budget: Examples beat Knobs1d◆Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs1d◆RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination1d◆HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization1d◆When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models1d◆SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation1d◆ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System1d◆LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets1d◆Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility1d◆Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation1d◆LIGO-PINN: Learned Initialization via Gated Optimization to Alleviate Convergence Failures in Physics Informed Neural Networks1d◆
News/Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity
arxiv
PublishedMay 8, 2026 at 4:00 AM
—neutral

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2605.06327v1 Announce Type: cross Abstract: Safety benchmarks are routinely treated as evidence about how a language model will behave once deployed, but this inference is fragile if behavior depends on whether a prompt looks like an evaluation. We define evaluation-context divergence as an ob

Models mentioned
02
  • 01meta-llama logo
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    DL 1.6M+0.7%IN $0.10/Mtok
  • 02meta-llama logo
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
Compare these 2 models→
Related
05
  • arxiv1d
    Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
  • arxiv3d
    A Shared Subcircuit Lets LLMs Count Down Across Tasks
  • techcrunch19d
    Vibe-coding platform Base44 launches own model as AI startups seek defensibility
  • arxivMay 28
    A Benchmark Construction and Evaluation Framework for Specialist Domains: Case Study on Defense-related Documents
  • arxivMay 22
    GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
07
  • 01
    OLMo-3
  • 02
    OLMo-3-Instruct
  • 03
    Mistral-Small-3.2
  • 04
    Phi-3.5-mini
  • 05
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    1.6M dl
  • 06
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
  • 07
    Llama-Guard-3-8B
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#safety#benchmark#evaluation#language models

No replies yet. Be first.

Mentioned models
07
  • 01
    OLMo-3
  • 02
    OLMo-3-Instruct
  • 03
    Mistral-Small-3.2
  • 04
    Phi-3.5-mini
  • 05
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    1.6M dl
  • 06
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
  • 07
    Llama-Guard-3-8B
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#safety#benchmark#evaluation#language models

Related coverage

More from ARXIV
arxivThe Steering Budget: Examples beat Knobs1darxivPolestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs1darxivRxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination1darxivHABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization1d
The Bubble Brief
WEEKLY

Read safety insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews