·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Architecting memory and storage in the AI era2h◆Roland is getting into generative AI music with Melody Flip3h◆What will Apple’s John Ternus era look like?4h◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge5h◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot5h◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack5h◆Google’s Gemini Spark can now manage your Google Photos library6h◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event7h◆Rogue OpenAI agents appear to have organized another attack using a German wiki7h◆Instagram’s AI detection is a mess (again)9h◆Why AI food looks like that10h◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers10h◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users10h◆This NAS company wants to run your local smart home11h◆Data from drones in Ukraine is fueling a new Wild West marketplace12h◆The sameness problem behind those unappetizing AI-generated menus17h◆CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning17h◆Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation17h◆X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System17h◆A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors17h◆Architecting memory and storage in the AI era2h◆Roland is getting into generative AI music with Melody Flip3h◆What will Apple’s John Ternus era look like?4h◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge5h◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot5h◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack5h◆Google’s Gemini Spark can now manage your Google Photos library6h◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event7h◆Rogue OpenAI agents appear to have organized another attack using a German wiki7h◆Instagram’s AI detection is a mess (again)9h◆Why AI food looks like that10h◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers10h◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users10h◆This NAS company wants to run your local smart home11h◆Data from drones in Ukraine is fueling a new Wild West marketplace12h◆The sameness problem behind those unappetizing AI-generated menus17h◆CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning17h◆Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation17h◆X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System17h◆A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors17h◆
News/WorkBench Revisited: Workplace Agents Two Years On
arxiv
PublishedJuly 2, 2026 at 4:00 AM
▲bullish

WorkBench Revisited: Workplace Agents Two Years On

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to date, Claude Fable 5, now completes 98%. Beyond this considerable progress in frontier agent perfor

Models mentioned
01
  • 01openai logo
    gpt-4
    openai/gpt-4
    0.0%IN $30.00/Mtok
Related
05
  • arxivJul 10
    A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents
  • arxivMay 19
    EmoMind: Decoding Affective Captions from Human Brain fMRI
  • arxivMay 11
    End-to-end PDDL Planning with Hardcoded and Dynamic Agents
  • arxivApr 28
    ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
  • arxivApr 21
    SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
02
  • 01
    gpt-4
    openai/gpt-4
  • 02
    Claude Fable 5
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#safety#open-source#performance
Mentioned companies
01
OpenAI

No replies yet. Be first.

Mentioned models
02
  • 01
    gpt-4
    openai/gpt-4
  • 02
    Claude Fable 5
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#safety#open-source#performance
Mentioned companies
01
OpenAI

Related coverage

More from ARXIV
arxivCulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning17harxivProactive Service Agents: A Unified Decision Framework, Methods, and Evaluation17harxivX-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System17harxivA Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors17h
The Bubble Brief
WEEKLY

Read benchmark insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews