arxiv
PublishedJuly 2, 2026 at 4:00 AM
▲bullish
WorkBench Revisited: Workplace Agents Two Years On
Publisher summary· verbatim
arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to date, Claude Fable 5, now completes 98%. Beyond this considerable progress in frontier agent perfor
Models mentioned
01Related
05- arxivJul 10A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents
- arxivMay 19EmoMind: Decoding Affective Captions from Human Brain fMRI
- arxivMay 11End-to-end PDDL Planning with Hardcoded and Dynamic Agents
- arxivApr 28ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
- arxivApr 21SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivCulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning17harxivProactive Service Agents: A Unified Decision Framework, Methods, and Evaluation17harxivX-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System17harxivA Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors17hThe Bubble Brief
WEEKLYRead benchmark insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗