arxiv
PublishedJuly 27, 2026 at 4:00 AM
—neutral
Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark
Publisher summary· verbatim
arXiv:2607.21685v1 Announce Type: new Abstract: A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings (MeSH), assigned either by
Models mentioned
01Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning6harxivSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks6harxivDistribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts6harxivIn RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning6hThe Bubble Brief
WEEKLYRead benchmark insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗