arxiv
PublishedSeptember 4, 2026 at 4:00 AM
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
Publisher summary· verbatim
arXiv:2609.04198v1 Announce Type: new Abstract: Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We aud
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivCulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning14harxivProactive Service Agents: A Unified Decision Framework, Methods, and Evaluation14harxivX-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System14harxivA Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors14hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗