arxiv
PublishedJuly 14, 2026 at 4:00 AM
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
Publisher summary· verbatim
arXiv:2507.11059v3 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset. Recent studies have uncovered severe data contamination is
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivADS-C: Antidistillation Sampling for Classification15harxivDigital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents15harxivEpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections15harxivBefore the Action: Benchmarking LLMs on Prospective Hypothesis Discovery15hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗