·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft5h◆Hikers rescued after using Google Gemini for planning8h◆OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure9h◆OpenAI admits to German wiki ‘incident’16h◆XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation1d◆OpenAI’s rogue agents keep escaping, with no formal process to investigate them1d◆AI compute provider Nscale is looking for $3.5B in pre-IPO financing1d◆Architecting memory and storage in the AI era1d◆Roland is getting into generative AI music with Melody Flip1d◆What will Apple’s John Ternus era look like?1d◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge1d◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot1d◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack1d◆Google’s Gemini Spark can now manage your Google Photos library1d◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event1d◆Rogue OpenAI agents appear to have organized another attack using a German wiki1d◆Instagram’s AI detection is a mess (again)1d◆Why AI food looks like that1d◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers1d◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users1d◆Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft5h◆Hikers rescued after using Google Gemini for planning8h◆OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure9h◆OpenAI admits to German wiki ‘incident’16h◆XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation1d◆OpenAI’s rogue agents keep escaping, with no formal process to investigate them1d◆AI compute provider Nscale is looking for $3.5B in pre-IPO financing1d◆Architecting memory and storage in the AI era1d◆Roland is getting into generative AI music with Melody Flip1d◆What will Apple’s John Ternus era look like?1d◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge1d◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot1d◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack1d◆Google’s Gemini Spark can now manage your Google Photos library1d◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event1d◆Rogue OpenAI agents appear to have organized another attack using a German wiki1d◆Instagram’s AI detection is a mess (again)1d◆Why AI food looks like that1d◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers1d◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users1d◆
News/model/Agents-A1

Agents-A1 news

50 articles mentioning Agents-A1

techcrunch1d ago

OpenAI’s rogue agents keep escaping, with no formal process to investigate them

OpenAI’s latest agent swarm incident adds urgency to calls for independent investigations as researchers and lawmakers question whether AI labs should control the scope of their own safety reviews.

techcrunch1d ago

Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge

It's the latest failure of OpenAI's internal monitoring and security systems.

theverge1d ago

Rogue OpenAI agents appear to have organized another attack using a German wiki

A swarm of rogue AI agents from OpenAI reportedly commandeered a German website and transformed it into a messaging board for other agents, with officials staying quiet about the incident for weeks as the company prepared to launch its most advanced model yet, Astra. The finding adds to intensifying

arxiv1d ago

Measuring Harmfulness of Computer-Using Agents

arXiv:2508.00935v3 Announce Type: replace-cross Abstract: Computer-using agents (CUAs), which can autonomously control computers to perform multi-step actions, might pose significant safety risks if misused. However, existing benchmarks mainly evaluate LMs in chatbots or simple tool use. To more com

arxiv1d ago

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

arXiv:2609.03588v1 Announce Type: new Abstract: As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability ac

arxiv1d ago

Evolving Excellence: Automated Optimization of LLM-based Agents

arXiv:2512.09108v2 Announce Type: replace-cross Abstract: Agentic AI systems built on large language models (LLMs) offer significant potential for automating complex workflows, from software development to customer support. However, LLM agents often underperform due to suboptimal configurations; poo

arxiv1d ago

Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

arXiv:2609.03727v1 Announce Type: new Abstract: Large language model agents can plan, invoke tools, and modify external states, yet most systems still take an explicit user instruction as a fixed starting point. Proactive service moves the decision upstream: an agent must infer service opportunities

arxiv1d ago

What Do CAE Simulation Agents Really Need Beyond a Generic Harness?

arXiv:2609.03718v1 Announce Type: cross Abstract: Computer-aided engineering (CAE) simulation is among the largest and most demanding areas of engineering, where setting up a solver such as OpenFOAM, FEniCS, or COMSOL takes real expertise. Large language model (LLM) agents promise to turn a natural-

arxiv1d ago

Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents

arXiv:2609.02889v1 Announce Type: new Abstract: A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strategy, format rules, and control heuristics. Existing reflective prompt-evolution m

arxiv1d ago

PatchBench: Evaluating AI Agents for Vulnerability Patching

arXiv:2609.04075v1 Announce Type: cross Abstract: AI agents have recently demonstrated strong performance in automated vulnerability patching. However, existing evaluations often validate a patch only by testing whether the provided Proof-of-Concept (PoC) input still triggers a crash. This leaves tw

arxiv1d ago

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

arXiv:2609.04167v1 Announce Type: cross Abstract: Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests and overlook review-derived acceptance constraints

arxiv1d ago

DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

arXiv:2609.03423v1 Announce Type: new Abstract: Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed

arxiv1d ago

From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

arXiv:2609.03590v1 Announce Type: cross Abstract: Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is essential for NextG interactive applications, yet providing strict End-to-End (E2E) peak latency guarantees remains an open challenge. Two obstacles limit the adop

arxiv1d ago

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

arXiv:2609.03467v1 Announce Type: cross Abstract: Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conve

arxiv1d ago

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

arXiv:2606.05553v2 Announce Type: replace-cross Abstract: Role-playing language agents (RPLAs) simulate specific characters and personas across applications such as entertainment, companionship, interactive storytelling, and education. Faithful role-play requires more than producing plausible, in-ch

arxiv1d ago

EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation

arXiv:2605.26785v2 Announce Type: replace-cross Abstract: Post-trained LLMs are often optimized to produce helpful, polite, and accommodating responses. In adversarial negotiation, however, such behavior can become a vulnerability: emotionally framed language may influence an agent's bargaining deci

arxiv1d ago

AI Agents Push Humans Out of the Loop

arXiv:2608.23642v2 Announce Type: replace Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current approaches to AI agent design impe

arxiv1d ago

RuleMem: Active Rule Memory for Long-Term Conversational Agents

arXiv:2609.03915v1 Announce Type: new Abstract: Question answering agents in long-term conversations must reason over massive, temporally dispersed dialogue histories. However, existing memory mechanisms primarily treat past information as \textit{passively} stored facts, leading to semantic gaps an

arxiv1d ago

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

arXiv:2608.22793v2 Announce Type: replace-cross Abstract: Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cannot yet, be s

arxiv1d ago

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

arXiv:2609.04159v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, and free-form genera

arxiv1d ago

The Natural Language Interaction Protocol and Standard for AI Agents

arXiv:2609.04135v1 Announce Type: new Abstract: AI agents are increasingly being developed and deployed across organizations using heterogeneous agent-development frameworks, AI models, tool interfaces, protocols, and execution environments. To realize their potential social and business impact, the

arxiv1d ago

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

arXiv:2609.03438v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to benign mistakes. A reliable agent should not only know how to act, but also

arxiv1d ago

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

arXiv:2608.27141v4 Announce Type: replace-cross Abstract: Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iteratio

arxiv1d ago

GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models

arXiv:2606.12821v2 Announce Type: replace Abstract: Environmental scientists spend disproportionate effort on data wrangling rather than analysis. New AI agents can be a helpful tool, but no benchmark exists to evaluate AI agents that automate environmental geospatial workflows through structured to

arxiv1d ago

LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents

arXiv:2606.18388v2 Announce Type: replace-cross Abstract: RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization parameters predominantly oscillate in response to shifting training d

arxiv1d ago

Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents

arXiv:2609.00065v2 Announce Type: replace-cross Abstract: A language-model agent asked to analyse an experiment will usually return working code. Whether the analysis is defensible is a different question. A defensible analysis depends on procedural choices: which test the field accepts, which ident

arxiv1d ago

CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents

arXiv:2606.02372v2 Announce Type: replace Abstract: Equipping language agents with world models enables them to anticipate environment dynamics and evaluate candidate actions before execution. However, existing textual world models are typically fixed after training, preventing them from adapting to

arxiv1d ago

RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

arXiv:2609.02902v1 Announce Type: new Abstract: Deploying task-oriented dialogue agents in enterprise customer support faces a persistent annotation bottleneck: robust training requires labelled interaction data at scale, yet enterprise conversational logs are privacy-sensitive and expensive to anno

arxiv1d ago

Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor

arXiv:2609.03221v1 Announce Type: new Abstract: Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor chang

arxiv1d ago

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

arXiv:2609.03383v1 Announce Type: new Abstract: Graph-based policy optimization improves credit assignment for long-horizon LLM agents by organizing rollout trajectories into state-transition graphs. However, existing methods construct graphs independently within each policy update, discarding trans

arxiv1d ago

Environment Evolution for Terminal Agents

arXiv:2609.04128v1 Announce Type: new Abstract: Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evoluti

arxiv1d ago

Speculative Macro Commit for Faster Tool-Using Agents

arXiv:2609.03236v1 Announce Type: new Abstract: Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and observation can delay subsequent decisions. We introduce \textbf{Speculative Macro C

arxiv2d ago

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

arXiv:2606.24026v2 Announce Type: replace Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize. In this work, we study whether language model (LM) ag

arxiv2d ago

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

arXiv:2606.17698v3 Announce Type: replace Abstract: As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose

arxiv2d ago

Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets

arXiv:2609.01918v1 Announce Type: new Abstract: Energy poverty is nearly absent from NLP-for-social-good, and the little existing work is either static retrieval/QA or relies on carbon-intensive cloud LLMs, a self-defeating "computational irony" for a humanitarian setting. We present EqGrid, a close

arxiv2d ago

Git4Data: Database-Native Version Control for AI Agents

arXiv:2609.02106v1 Announce Type: cross Abstract: Large Language Model (LLM) agents increasingly explore many candidate states of relational data in parallel, each of which should remain isolated, reproducible, and auditable, preferably through the same SQL interface used for ordinary data work. Exi

arxiv2d ago

Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

arXiv:2609.02760v1 Announce Type: new Abstract: On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no

arxiv2d ago

AI agents reshape consensus formation in human groups

arXiv:2609.02122v1 Announce Type: new Abstract: As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative des

arxiv2d ago

PGMem: Tightly Coupled Persona-Memory Graph for Lifelong Personalized Agents

arXiv:2608.01708v2 Announce Type: replace Abstract: Long-term personalized dialogue agents must track user preferences as their personas evolve. Existing memory systems organize past events well, but store personas as flat profiles detached from the events that justify them. This loose coupling lead

arxiv2d ago

WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling

arXiv:2609.01608v1 Announce Type: cross Abstract: Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces. Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation or trial-and-error

arxiv2d ago

Agents That Model Agents: Five Principles Toward a Theory of Mind for 6G Networks

arXiv:2609.01779v1 Announce Type: cross Abstract: Future 6G networks will rely on Large Language Model (LLM) agents to manage the Radio Access Network (RAN). However, current architectures assume inter-agent messages convey objective facts. A message is instead a \emph{trace} of the sender's reasoni

arxiv2d ago

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

arXiv:2609.02265v1 Announce Type: new Abstract: Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary

arxiv2d ago

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

arXiv:2604.11557v3 Announce Type: replace Abstract: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However, existing research exhibits inconsistent interaction representations, largely overlooks the str

arxiv2d ago

When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

arXiv:2609.01985v1 Announce Type: new Abstract: As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization of how they behave on systems-level requirements: schema design, async orchestration, configuration correctness, and retrieval-filtering trade-off

arxiv2d ago

Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence

arXiv:2609.01873v1 Announce Type: new Abstract: Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly

arxiv2d ago

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

arXiv:2609.01660v1 Announce Type: cross Abstract: Production deployments of large language model (LLM) agents remain unreliable on long, multi-step workflows even as benchmark success rates climb steadily. We argue this gap is largely an artifact of task horizon: benchmarks are dominated by short-to

arxiv2d ago

Can Coding Agents Reproduce Findings in Computational Materials Science?

arXiv:2605.00803v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific

arxiv2d ago

AdaMem: Learning What to Remember with Adaptive Memory Policies for Personalized Agents

arXiv:2606.21144v2 Announce Type: replace-cross Abstract: Long-term memory systems allow LLM agents to preserve information beyond a single context window, but most systems focus on storing and retrieving facts after extraction, leaving the write decision under-specified. What deserves memory can de

arxiv2d ago

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

arXiv:2609.01852v1 Announce Type: new Abstract: Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capability changes. We evaluate a frozen, closed-set, action-scored benchmark w

arxiv2d ago

Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision

arXiv:2609.02057v1 Announce Type: new Abstract: Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits are unavailable. In this work, we study prefix-level risk prediction for web agents using observable trajectory signals: given an evolving prefix, e

HomeModelsNews