·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning6h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks6h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts6h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning6h◆FrontierChallenge: Evaluating Scientific Workflow Completion6h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier6h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising6h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic6h◆Omni Interaction Agent Technical Report6h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification6h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability6h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization6h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding6h◆Tracing Computation Density in LLMs6h◆Cultural Binding Heads in Language Models6h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training6h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models6h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning6h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection6h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation6h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning6h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks6h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts6h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning6h◆FrontierChallenge: Evaluating Scientific Workflow Completion6h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier6h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising6h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic6h◆Omni Interaction Agent Technical Report6h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification6h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability6h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization6h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding6h◆Tracing Computation Density in LLMs6h◆Cultural Binding Heads in Language Models6h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training6h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models6h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning6h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection6h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation6h◆
News/Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark
arxiv
PublishedJuly 27, 2026 at 4:00 AM
—neutral

Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2607.21685v1 Announce Type: new Abstract: A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings (MeSH), assigned either by

Models mentioned
01
  • 01dmis-lab logo
    biobert-v1.1
    dmis-lab/biobert-v1.1
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
02
  • 01
    biobert-v1.1
    dmis-lab/biobert-v1.1
  • 02
    bag-of-words logistic regression
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#evaluation#classification#nlp

No replies yet. Be first.

Mentioned models
02
  • 01
    biobert-v1.1
    dmis-lab/biobert-v1.1
  • 02
    bag-of-words logistic regression
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#evaluation#classification#nlp

Related coverage

More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning6harxivSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks6harxivDistribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts6harxivIn RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning6h
The Bubble Brief
WEEKLY

Read benchmark insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews