·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
How Deutsche Telekom is rewiring telecommunications with AI6h◆When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals9h◆PLURAL: A Global Dataset for Value Alignment9h◆Does Dimensionality Reduction via Random Projections Preserve Landscape Features?9h◆Infinity-Parser2 Technical Report9h◆Agentic Neural Architecture Search9h◆Leveraging Color Naming for Image Enhancement9h◆Svarna: An Open Corpus Workbench for Modern Greek9h◆Distributed Sketching on Data Partitions for OLS Regression9h◆TTHE: Test-Time Harness Evolution9h◆LTM: Large-scale Terrain Model for Wildfire-prone Landscapes9h◆($\theta_l, \theta_u$)-Parametric Multi-Task Optimization: Joint Search in Solution and Infinite Task Spaces9h◆Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention9h◆Aleena: Alignment Agent for Research Software Engineering Collaborations9h◆Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs9h◆MetaHGNIE: Meta-Path Induced Hypergraph Contrastive Learning in Heterogeneous Knowledge Graphs9h◆MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction9h◆GradInf: Gradient Estimation as Probabilistic Inference9h◆Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring9h◆Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning9h◆How Deutsche Telekom is rewiring telecommunications with AI6h◆When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals9h◆PLURAL: A Global Dataset for Value Alignment9h◆Does Dimensionality Reduction via Random Projections Preserve Landscape Features?9h◆Infinity-Parser2 Technical Report9h◆Agentic Neural Architecture Search9h◆Leveraging Color Naming for Image Enhancement9h◆Svarna: An Open Corpus Workbench for Modern Greek9h◆Distributed Sketching on Data Partitions for OLS Regression9h◆TTHE: Test-Time Harness Evolution9h◆LTM: Large-scale Terrain Model for Wildfire-prone Landscapes9h◆($\theta_l, \theta_u$)-Parametric Multi-Task Optimization: Joint Search in Solution and Infinite Task Spaces9h◆Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention9h◆Aleena: Alignment Agent for Research Software Engineering Collaborations9h◆Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs9h◆MetaHGNIE: Meta-Path Induced Hypergraph Contrastive Learning in Heterogeneous Knowledge Graphs9h◆MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction9h◆GradInf: Gradient Estimation as Probabilistic Inference9h◆Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring9h◆Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning9h◆
News/RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
arxiv
PublishedJune 20, 2026 at 4:00 AM
—neutral

RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.15862v2 Announce Type: replace Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simula

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#evaluation#autonomy#decision-making

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#evaluation#autonomy#decision-making

Related coverage

More from ARXIV
arxivWhen LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals9harxivPLURAL: A Global Dataset for Value Alignment9harxivDoes Dimensionality Reduction via Random Projections Preserve Landscape Features?9harxivInfinity-Parser2 Technical Report9h
The Bubble Brief
WEEKLY

Read benchmark insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews