·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents56m◆AI and the rise of the universal entertainment app1h◆Substack adds an AI detector to help spot blogs written by no one1h◆Data centers expected to use 4x more electricity by 20352h◆Google releases three new Gemini models — but no 3.5 Pro3h◆Introducing the ChatGPT for small business program3h◆Anthropic’s $1.5 billion book piracy settlement approved by judge3h◆US threatens sanctions against Chinese AI models over IP theft5h◆Google launches a cheaper alternative to large AI security models like Mythos5h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated7h◆Halliday’s latest smart glasses feature a much-improved display7h◆America needs to stop getting shocked by Chinese AI9h◆Advancing next-gen AI with materials science innovation10h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else10h◆Capacity and Redundancy Trade-offs in Multi-Task Learning16h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation16h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making16h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection16h◆Supervised Reward Inference16h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization16h◆Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents56m◆AI and the rise of the universal entertainment app1h◆Substack adds an AI detector to help spot blogs written by no one1h◆Data centers expected to use 4x more electricity by 20352h◆Google releases three new Gemini models — but no 3.5 Pro3h◆Introducing the ChatGPT for small business program3h◆Anthropic’s $1.5 billion book piracy settlement approved by judge3h◆US threatens sanctions against Chinese AI models over IP theft5h◆Google launches a cheaper alternative to large AI security models like Mythos5h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated7h◆Halliday’s latest smart glasses feature a much-improved display7h◆America needs to stop getting shocked by Chinese AI9h◆Advancing next-gen AI with materials science innovation10h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else10h◆Capacity and Redundancy Trade-offs in Multi-Task Learning16h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation16h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making16h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection16h◆Supervised Reward Inference16h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization16h◆
Tag

#software engineering

6 articles tagged #software engineering

arxivJul 1bullish

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

arXiv:2606.31551v1 Announce Type: new Abstract: Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software engineering and other long-horizon tasks. A central challenge is that autonomous post-training is no

GPDE2 models#autonomous training#language models#benchmarkRead on arxiv →
arxivJun 12bullish

On Sequence-to-Sequence Models for Automated Log Parsing

arXiv:2602.07698v2 Announce Type: replace-cross Abstract: Context: Log parsing is a critical standard operating procedure in software systems, enabling monitoring, anomaly detection, and failure diagnosis. However, automated log parsing remains challenging due to heterogeneous log formats, distribut

TRMALS5 models · +2#log parsing#sequence modelling#software engineeringRead on arxiv →
arxivMay 11

The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking

arXiv:2605.06707v1 Announce Type: cross Abstract: This paper presents an eight-week observational comparison of 68 single-file HTML generations collected across 17 public experiments in the "HTML AI Battle" project between December 10, 2025 and February 4, 2026. Four reasoning model families, GPT, G

GPGEGR4 models · +1#software engineering#artificial intelligence#benchmarkRead on arxiv →
arxivMay 1bearish

Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows

arXiv:2604.28139v1 Announce Type: cross Abstract: LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at release time and grade mainly the final response, making it difficult

#benchmark#workflow#evaluationRead on arxiv →
arxivApr 24bullish

DryRUN: On the Role of Public Tests in LLM-Driven Code Generation

arXiv:2604.21598v1 Announce Type: cross Abstract: Multi-agent frameworks are widely used in autonomous code generation and have applications in complex algorithmic problem-solving. Recent work has addressed the challenge of generating functionally correct code by incorporating simulation-driven plan

DRCO2 models#autonomous code generation#software engineering#large language modelsRead on arxiv →
arxivApr 16bullish

AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection

arXiv:2604.11950v1 Announce Type: cross Abstract: While recent LLM-based agents can identify many candidate bugs in source code, their reports remain static hypotheses that require manual validation, limiting the practicality of automated bug detection. We frame this challenge as a test generation t

CLCO2 models#software engineering#bug detection#test generationRead on arxiv →
HomeModelsNews