·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
SPINE: Bridging the Cyber-Physical Gap with Agentic AI21m◆AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation21m◆Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management21m◆Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable21m◆Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases21m◆EZSMT Version 3, Matured21m◆Set-shifting Behavioral Test for Harnessed Agents21m◆LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning21m◆How Far Can Root Cause Analysis Go on Real-World Telemetry Data?21m◆Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling21m◆AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized21m◆SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing21m◆Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System21m◆STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle21m◆UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following21m◆Explaining Reinforcement Learning Agents via Inductive Logic Programming21m◆When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects21m◆AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities21m◆CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems21m◆Experience Memory Graph: One-Shot Error Correction for Agents21m◆SPINE: Bridging the Cyber-Physical Gap with Agentic AI21m◆AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation21m◆Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management21m◆Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable21m◆Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases21m◆EZSMT Version 3, Matured21m◆Set-shifting Behavioral Test for Harnessed Agents21m◆LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning21m◆How Far Can Root Cause Analysis Go on Real-World Telemetry Data?21m◆Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling21m◆AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized21m◆SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing21m◆Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System21m◆STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle21m◆UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following21m◆Explaining Reinforcement Learning Agents via Inductive Logic Programming21m◆When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects21m◆AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities21m◆CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems21m◆Experience Memory Graph: One-Shot Error Correction for Agents21m◆
News/WorkBench Revisited: Workplace Agents Two Years On
arxiv
PublishedJuly 2, 2026 at 4:00 AM
▲bullish

WorkBench Revisited: Workplace Agents Two Years On

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to date, Claude Fable 5, now completes 98%. Beyond this considerable progress in frontier agent perfor

Models mentioned
01
  • 01openai logo
    gpt-4
    openai/gpt-4
    0.0%IN $30.00/Mtok
Related
05
  • arxiv6d
    A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents
  • arxivMay 19
    EmoMind: Decoding Affective Captions from Human Brain fMRI
  • arxivMay 11
    End-to-end PDDL Planning with Hardcoded and Dynamic Agents
  • arxivApr 28
    ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
  • arxivApr 21
    SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
02
  • 01
    gpt-4
    openai/gpt-4
  • 02
    Claude Fable 5
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#safety#open-source#performance
Mentioned companies
01
OpenAI

No replies yet. Be first.

Mentioned models
02
  • 01
    gpt-4
    openai/gpt-4
  • 02
    Claude Fable 5
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#safety#open-source#performance
Mentioned companies
01
OpenAI

Related coverage

More from ARXIV
arxivSPINE: Bridging the Cyber-Physical Gap with Agentic AI21marxivAI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation21marxivCost-Optimal Foundation Model Deployment Portfolio for Transportation Management21marxivHarness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable21m
The Bubble Brief
WEEKLY

Read benchmark insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews