·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?2h◆A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution2h◆Looped State-Space Language Models with Adaptive Exit-State Selection2h◆ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory2h◆Laguerre Geometry for Interpreting Large Language Models2h◆Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification2h◆MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis2h◆Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories2h◆WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs2h◆Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF2h◆Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems2h◆Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning2h◆Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health2h◆First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint)2h◆SETA: Scaling Environments for Terminal Agents2h◆Incremental Transformer for Surrogate-Based Inverse Design of Geopolymer Mixtures2h◆Learning Linear Temporal Specifications from Demonstrations with Uncertainty2h◆STAMP: Provenance-Guided Credit Assignment for Deep Search Agents2h◆The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy2h◆SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL2h◆Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?2h◆A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution2h◆Looped State-Space Language Models with Adaptive Exit-State Selection2h◆ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory2h◆Laguerre Geometry for Interpreting Large Language Models2h◆Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification2h◆MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis2h◆Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories2h◆WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs2h◆Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF2h◆Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems2h◆Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning2h◆Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health2h◆First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint)2h◆SETA: Scaling Environments for Terminal Agents2h◆Incremental Transformer for Surrogate-Based Inverse Design of Geopolymer Mixtures2h◆Learning Linear Temporal Specifications from Demonstrations with Uncertainty2h◆STAMP: Provenance-Guided Credit Assignment for Deep Search Agents2h◆The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy2h◆SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL2h◆
News/How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent
arxiv
PublishedApril 29, 2026 at 4:00 AM

How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2604.07236v4 Announce Type: replace Abstract: Agent harnesses -- the stateful programs that wrap a language model and decide what it sees at each step -- are now known to change end-to-end performance on a fixed model by as much as six times. That raises a question asked less often than it sho

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivWho&When Pro: Can LLMs Really Attribute Failures in AI Agents?2harxivA Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution2harxivLooped State-Space Language Models with Adaptive Exit-State Selection2harxivABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory2h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews