·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance7m◆Boosting Adversarial Robustness and Generalization with Dictionary Structure7m◆A Dominant Supplier Slows Recursive Drift More Than It Steers It7m◆COMiT: Learning Structured Visual Tokens through Sequential Communication7m◆Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation7m◆Data-Free Pruning of Self-Attention Layers in LLMs7m◆Screening Is Enough7m◆Data Unlearning via Inverse Distillation7m◆ReCIRC: Rectified Conformal Risk Control7m◆Replay-buffer engineering for noise-aware quantum circuit optimization7m◆SINO: Scale-Invariant Neural Operator7m◆Preferent Compression Bounds Are Tight7m◆Fundamental Limits of Transferability and Equivariance in Algebraic Signal Models I: Finite Dimensions7m◆How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models7m◆Bregman Consensus7m◆In-Context Learning for Robots: Methods and Applications7m◆Modal Logic Neural Networks7m◆Volatility-Clustering Adaptation for Financial Time Series7m◆Reference-Guided Machine Unlearning7m◆Agentic Federated Learning: Rule-Based Client and Server Agents for Adaptive Training7m◆Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance7m◆Boosting Adversarial Robustness and Generalization with Dictionary Structure7m◆A Dominant Supplier Slows Recursive Drift More Than It Steers It7m◆COMiT: Learning Structured Visual Tokens through Sequential Communication7m◆Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation7m◆Data-Free Pruning of Self-Attention Layers in LLMs7m◆Screening Is Enough7m◆Data Unlearning via Inverse Distillation7m◆ReCIRC: Rectified Conformal Risk Control7m◆Replay-buffer engineering for noise-aware quantum circuit optimization7m◆SINO: Scale-Invariant Neural Operator7m◆Preferent Compression Bounds Are Tight7m◆Fundamental Limits of Transferability and Equivariance in Algebraic Signal Models I: Finite Dimensions7m◆How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models7m◆Bregman Consensus7m◆In-Context Learning for Robots: Methods and Applications7m◆Modal Logic Neural Networks7m◆Volatility-Clustering Adaptation for Financial Time Series7m◆Reference-Guided Machine Unlearning7m◆Agentic Federated Learning: Rule-Based Client and Server Agents for Adaptive Training7m◆
News/Benchy: towards a universal language for task-oriented AI benchmarks
arxiv
PublishedSeptember 29, 2026 at 4:00 AM
—neutral

Benchy: towards a universal language for task-oriented AI benchmarks

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2609.30550v1 Announce Type: new Abstract: Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI)

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivPredictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance7marxivBoosting Adversarial Robustness and Generalization with Dictionary Structure7marxivA Dominant Supplier Slows Recursive Drift More Than It Steers It7marxivCOMiT: Learning Structured Visual Tokens through Sequential Communication7m
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
Built by Marouane Gazouzi
HomeModelsNews