·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
OpenAI admits to German wiki ‘incident’3h◆XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation15h◆OpenAI’s rogue agents keep escaping, with no formal process to investigate them15h◆AI compute provider Nscale is looking for $3.5B in pre-IPO financing17h◆Architecting memory and storage in the AI era19h◆Roland is getting into generative AI music with Melody Flip20h◆What will Apple’s John Ternus era look like?21h◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge22h◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot22h◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack22h◆Google’s Gemini Spark can now manage your Google Photos library23h◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event1d◆Rogue OpenAI agents appear to have organized another attack using a German wiki1d◆Instagram’s AI detection is a mess (again)1d◆Why AI food looks like that1d◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers1d◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users1d◆This NAS company wants to run your local smart home1d◆Data from drones in Ukraine is fueling a new Wild West marketplace1d◆The sameness problem behind those unappetizing AI-generated menus1d◆OpenAI admits to German wiki ‘incident’3h◆XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation15h◆OpenAI’s rogue agents keep escaping, with no formal process to investigate them15h◆AI compute provider Nscale is looking for $3.5B in pre-IPO financing17h◆Architecting memory and storage in the AI era19h◆Roland is getting into generative AI music with Melody Flip20h◆What will Apple’s John Ternus era look like?21h◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge22h◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot22h◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack22h◆Google’s Gemini Spark can now manage your Google Photos library23h◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event1d◆Rogue OpenAI agents appear to have organized another attack using a German wiki1d◆Instagram’s AI detection is a mess (again)1d◆Why AI food looks like that1d◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers1d◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users1d◆This NAS company wants to run your local smart home1d◆Data from drones in Ukraine is fueling a new Wild West marketplace1d◆The sameness problem behind those unappetizing AI-generated menus1d◆
News/AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
arxiv
PublishedApril 6, 2026 at 4:00 AM
▼bearish

AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2604.02947v1 Announce Type: new Abstract: Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across interactions and translate intermediate outputs into concrete actions. T

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
07
  • 01
    Claude Code
  • 02
    OpenClaw
  • 03
    IFlow
  • 04
    Qwen3-Coder
  • 05
    Kimi
  • 06
    GLM
  • 07
    DeepSeek
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#safety#benchmark#autonomous agents#language models

No replies yet. Be first.

Mentioned models
07
  • 01
    Claude Code
  • 02
    OpenClaw
  • 03
    IFlow
  • 04
    Qwen3-Coder
  • 05
    Kimi
  • 06
    GLM
  • 07
    DeepSeek
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#safety#benchmark#autonomous agents#language models
The Bubble Brief
WEEKLY

Read safety insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews