·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
China delivers a one-two punch to America’s AI dominance4h◆AI is more likely than humans to form biases when hiring6h◆Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal11h◆SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents11h◆PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment11h◆Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length11h◆Label-Free Concept Drift Assessment for Reliable AI in Emerging Wireless Applications11h◆DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales11h◆Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors11h◆CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data11h◆cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data11h◆Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?11h◆Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes11h◆From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems11h◆Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI11h◆The AI Fiction Paradox11h◆Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents11h◆EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections11h◆Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery11h◆How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI11h◆China delivers a one-two punch to America’s AI dominance4h◆AI is more likely than humans to form biases when hiring6h◆Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal11h◆SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents11h◆PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment11h◆Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length11h◆Label-Free Concept Drift Assessment for Reliable AI in Emerging Wireless Applications11h◆DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales11h◆Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors11h◆CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data11h◆cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data11h◆Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?11h◆Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes11h◆From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems11h◆Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI11h◆The AI Fiction Paradox11h◆Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents11h◆EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections11h◆Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery11h◆How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI11h◆
News/Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
arxiv
PublishedJuly 18, 2026 at 4:00 AM
—neutral

Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2607.14108v1 Announce Type: cross Abstract: This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory. To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivBeyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal11harxivSkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents11harxivPolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment11harxivLatency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length11h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews