·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
OpenAI bets on families as ChatGPT goes deeper into households11h◆Svarna: An Open Corpus Workbench for Modern Greek21h◆PLURAL: A Global Dataset for Value Alignment21h◆Validating LLMs in social science: Epistemic threats and emerging norms21h◆How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism21h◆WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search21h◆Temporal Preference Concepts and their Functions in a Large Language Model21h◆UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks21h◆When Synthetic Speech Is All You Have: Better Call GRPO21h◆ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation21h◆Towards Isolated Interventions via Almost Orthogonal Features in Language Models21h◆MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction21h◆Peer-Predictive Self-Training for Language Model Reasoning21h◆DeepTutor: Towards Agentic Personalized Tutoring21h◆COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation21h◆Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator21h◆CausalDS: Benchmarking Causal Reasoning in Data-Science Agents21h◆Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization21h◆Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition21h◆Where do LLMs Fall Short in CBT-Guided Affective Reasoning?21h◆OpenAI bets on families as ChatGPT goes deeper into households11h◆Svarna: An Open Corpus Workbench for Modern Greek21h◆PLURAL: A Global Dataset for Value Alignment21h◆Validating LLMs in social science: Epistemic threats and emerging norms21h◆How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism21h◆WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search21h◆Temporal Preference Concepts and their Functions in a Large Language Model21h◆UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks21h◆When Synthetic Speech Is All You Have: Better Call GRPO21h◆ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation21h◆Towards Isolated Interventions via Almost Orthogonal Features in Language Models21h◆MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction21h◆Peer-Predictive Self-Training for Language Model Reasoning21h◆DeepTutor: Towards Agentic Personalized Tutoring21h◆COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation21h◆Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator21h◆CausalDS: Benchmarking Causal Reasoning in Data-Science Agents21h◆Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization21h◆Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition21h◆Where do LLMs Fall Short in CBT-Guided Affective Reasoning?21h◆
News/Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving
arxiv
PublishedJune 30, 2026 at 4:00 AM
—neutral

Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.29493v1 Announce Type: new Abstract: Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof. However, the kernel only checks that a proof establishes a \emph{formal} statement; it doe

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivSvarna: An Open Corpus Workbench for Modern Greek21harxivPLURAL: A Global Dataset for Value Alignment21harxivValidating LLMs in social science: Epistemic threats and emerging norms21harxivHow Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism21h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews