·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning8h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks8h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts8h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning8h◆FrontierChallenge: Evaluating Scientific Workflow Completion8h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier8h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising8h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic8h◆Omni Interaction Agent Technical Report8h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification8h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability8h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization8h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding8h◆Tracing Computation Density in LLMs8h◆Cultural Binding Heads in Language Models8h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training8h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models8h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning8h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection8h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation8h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning8h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks8h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts8h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning8h◆FrontierChallenge: Evaluating Scientific Workflow Completion8h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier8h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising8h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic8h◆Omni Interaction Agent Technical Report8h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification8h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability8h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization8h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding8h◆Tracing Computation Density in LLMs8h◆Cultural Binding Heads in Language Models8h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training8h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models8h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning8h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection8h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation8h◆
Tag

#mathematics

11 articles tagged #mathematics

arxivJul 24bullish

LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization

arXiv:2607.20503v1 Announce Type: new Abstract: We present and evaluate LeanFlow, an LLM agent system specialized for translating mathematical papers into buildable Lean projects. Recent verifier-in-the-loop systems show that large formal artifacts can be produced, but it remains unclear which runti

KIGP2 models#mathematics#formalization#large language models
HomeModelsNews
Read on arxiv →
arxivJul 1bullish

Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics

arXiv:2606.31134v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evade human detection. Formal mathematical languages like Lean 4 offer mechanical proof checking, strong

LA1 model#autoformalization#mathematics#proof-checkingRead on arxiv →
arxivJun 25bearish

Riemann-Bench: A Benchmark for Moonshot Mathematics

arXiv:2604.06802v3 Announce Type: replace Abstract: Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competition-style problem solving. However, competition mathematics represents only a narrow slice of m

#mathematics#benchmark#researchRead on arxiv →
arxivJun 18bullish

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

arXiv:2505.23851v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization with genuine reasoning. To address this gap, we present ASyMOB, a high-resolution dataset of 35,368 va

#mathematics#evaluation#benchmarkRead on arxiv →
arxivJun 16bullish

SorryDB: Can AI Provers Complete Real-World Lean Theorems?

arXiv:2603.02668v2 Announce Type: replace Abstract: We present SorryDB, a dynamically-updating benchmark of open Lean tasks drawn from 78 real world formalization projects on GitHub. Unlike existing static benchmarks, often composed of competition problems, hillclimbing the SorryDB benchmark will yi

GE1 model#benchmark#open-source#mathematicsRead on arxiv →
arxivJun 12

Minimal surfaces, Knots, and Neural Networks

arXiv:2605.26234v2 Announce Type: replace-cross Abstract: A recent conjecture by Joel Fine posits a relationship between the coefficients of the HOMFLY polynomial of a knot $K$ in the 3-sphere $S^3$, and the signed count of minimal surfaces in hyperbolic 4-space $\mathrm{H}^4$ meeting the sphere at

PH1 model#mathematics#machine-learning#geometryRead on arxiv →
arxivJun 10bearish

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

arXiv:2606.10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-examined. To bridge this gap, we

LA1 model#evaluation#benchmark#mathematicsRead on arxiv →
arxivMay 29bullish

Formalizing Mathematics at Scale

arXiv:2605.29955v1 Announce Type: new Abstract: We present AutoformBot, a multi-agent system for building an Autoformalized Textbook Library At Scale (Atlas) in Lean 4. AutoformBot orchestrates thousands of LLM agents, equipped with formal verification tools, dependency-aware task scheduling, and co

AU1 model#autoformalization#mathematics#verificationRead on arxiv →
arxivMay 16

Numerical exploration of the range of shape functionals using neural networks

arXiv:2602.14881v2 Announce Type: replace-cross Abstract: We introduce a novel numerical framework for the exploration of Blaschke--Santal\'o diagrams, which are efficient tools characterizing the possible inequalities relating some given shape functionals. We introduce a parametrization of convex b

IN1 model#optimization#neural-networks#geometryRead on arxiv →
arxivMay 1bullish

QED: An Open-Source Multi-Agent System for Generating Mathematical Proofs on Open Problems

arXiv:2604.24021v2 Announce Type: replace Abstract: We explore a central question in AI for mathematics: can AI systems produce original, nontrivial proofs for open research problems? Despite strong benchmark performance, producing genuinely novel proofs remains an outstanding challenge for LLMs. Th

LLQE2 models#proof-generation#open-source#mathematicsRead on arxiv →
arxivApr 24

MathDuels: Evaluating LLMs as Problem Posers and Solvers

arXiv:2604.21916v1 Announce Type: new Abstract: As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, largely because they cast models solely as solvers of fixed problem sets. W

#benchmark#evaluation#language-modelsRead on arxiv →