·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Musk’s faster path to more gas turbines comes with pollution problem7h◆Texas Governor Abbott blocks funding for more Flock cameras8h◆Caterpillar is bringing to AI deployment what it learned from automating mining9h◆Sony Music, Warner sue Anthropic, alleging a “brazen campaign” of intellectual property theft1d◆Sony Music Publishing and Warner Chappell are suing Anthropic1d◆“We’re not doing 30 bets a year”: Vijay Pande on betting small after running $4 billion at a16z1d◆Nvidia’s AI advantage is moving beyond the GPU1d◆Musicians-turned-detectives are hunting for AI grifters1d◆Fine-Tuning of Transformer models with Frames1d◆Feature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection1d◆Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search1d◆Syntax vs. Semantics: How Transformers Learn Deep Dependencies1d◆NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation1d◆FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets1d◆Unsupervised Post-Training of Foundation Models: A Survey1d◆Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating1d◆Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement1d◆Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments1d◆Magnon-induced phononic Chern insulator1d◆Five Primitives for Governing Autonomous AI Agents at Runtime1d◆Musk’s faster path to more gas turbines comes with pollution problem7h◆Texas Governor Abbott blocks funding for more Flock cameras8h◆Caterpillar is bringing to AI deployment what it learned from automating mining9h◆Sony Music, Warner sue Anthropic, alleging a “brazen campaign” of intellectual property theft1d◆Sony Music Publishing and Warner Chappell are suing Anthropic1d◆“We’re not doing 30 bets a year”: Vijay Pande on betting small after running $4 billion at a16z1d◆Nvidia’s AI advantage is moving beyond the GPU1d◆Musicians-turned-detectives are hunting for AI grifters1d◆Fine-Tuning of Transformer models with Frames1d◆Feature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection1d◆Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search1d◆Syntax vs. Semantics: How Transformers Learn Deep Dependencies1d◆NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation1d◆FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets1d◆Unsupervised Post-Training of Foundation Models: A Survey1d◆Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating1d◆Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement1d◆Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments1d◆Magnon-induced phononic Chern insulator1d◆Five Primitives for Governing Autonomous AI Agents at Runtime1d◆
News/Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity
arxiv
PublishedMay 8, 2026 at 4:00 AM
—neutral

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2605.06327v1 Announce Type: cross Abstract: Safety benchmarks are routinely treated as evidence about how a language model will behave once deployed, but this inference is fragile if behavior depends on whether a prompt looks like an evaluation. We define evaluation-context divergence as an ob

Models mentioned
02
  • 01meta-llama logo
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    DL 1.6M0.0%IN $0.10/Mtok
  • 02meta-llama logo
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
Compare these 2 models→
Related
05
  • arxivJul 29
    FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
  • arxivJul 18
    Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
  • arxivJul 16
    A Shared Subcircuit Lets LLMs Count Down Across Tasks
  • techcrunchJun 30
    Vibe-coding platform Base44 launches own model as AI startups seek defensibility
  • arxivMay 28
    A Benchmark Construction and Evaluation Framework for Specialist Domains: Case Study on Defense-related Documents
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
07
  • 01
    OLMo-3
  • 02
    OLMo-3-Instruct
  • 03
    Mistral-Small-3.2
  • 04
    Phi-3.5-mini
  • 05
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    1.6M dl
  • 06
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
  • 07
    Llama-Guard-3-8B
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#safety#benchmark#evaluation#language models

No replies yet. Be first.

Mentioned models
07
  • 01
    OLMo-3
  • 02
    OLMo-3-Instruct
  • 03
    Mistral-Small-3.2
  • 04
    Phi-3.5-mini
  • 05
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    1.6M dl
  • 06
    Llama-3.1-70B
    meta-llama/Llama-3.1-70B
  • 07
    Llama-Guard-3-8B
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#safety#benchmark#evaluation#language models

Related coverage

More from ARXIV
arxivFine-Tuning of Transformer models with Frames1darxivFeature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection1darxivNaive Prompt Optimization: Rethinking the Need for Complex Prompt Search1darxivSyntax vs. Semantics: How Transformers Learn Deep Dependencies1d
The Bubble Brief
WEEKLY

Read safety insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews