·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Svarna: An Open Corpus Workbench for Modern Greek5h◆PLURAL: A Global Dataset for Value Alignment5h◆Validating LLMs in social science: Epistemic threats and emerging norms5h◆How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism5h◆WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search5h◆Temporal Preference Concepts and their Functions in a Large Language Model5h◆UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks5h◆When Synthetic Speech Is All You Have: Better Call GRPO5h◆ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation5h◆Towards Isolated Interventions via Almost Orthogonal Features in Language Models5h◆MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction5h◆Peer-Predictive Self-Training for Language Model Reasoning5h◆DeepTutor: Towards Agentic Personalized Tutoring5h◆COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation5h◆Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator5h◆CausalDS: Benchmarking Causal Reasoning in Data-Science Agents5h◆Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization5h◆Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition5h◆Where do LLMs Fall Short in CBT-Guided Affective Reasoning?5h◆Holographic Neural PCFG for Unsupervised Parsing5h◆Svarna: An Open Corpus Workbench for Modern Greek5h◆PLURAL: A Global Dataset for Value Alignment5h◆Validating LLMs in social science: Epistemic threats and emerging norms5h◆How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism5h◆WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search5h◆Temporal Preference Concepts and their Functions in a Large Language Model5h◆UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks5h◆When Synthetic Speech Is All You Have: Better Call GRPO5h◆ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation5h◆Towards Isolated Interventions via Almost Orthogonal Features in Language Models5h◆MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction5h◆Peer-Predictive Self-Training for Language Model Reasoning5h◆DeepTutor: Towards Agentic Personalized Tutoring5h◆COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation5h◆Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator5h◆CausalDS: Benchmarking Causal Reasoning in Data-Science Agents5h◆Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization5h◆Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition5h◆Where do LLMs Fall Short in CBT-Guided Affective Reasoning?5h◆Holographic Neural PCFG for Unsupervised Parsing5h◆
News/The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
arxiv
PublishedMay 11, 2026 at 4:00 AM
—neutral

The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2605.06707v1 Announce Type: cross Abstract: This paper presents an eight-week observational comparison of 68 single-file HTML generations collected across 17 public experiments in the "HTML AI Battle" project between December 10, 2025 and February 4, 2026. Four reasoning model families, GPT, G

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
04
  • 01
    GPT
  • 02
    Gemini
  • 03
    Grok
  • 04
    Claude
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#software engineering#artificial intelligence#benchmark#evaluation

No replies yet. Be first.

Mentioned models
04
  • 01
    GPT
  • 02
    Gemini
  • 03
    Grok
  • 04
    Claude
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#software engineering#artificial intelligence#benchmark#evaluation

Related coverage

More from ARXIV
arxivSvarna: An Open Corpus Workbench for Modern Greek5harxivPLURAL: A Global Dataset for Value Alignment5harxivValidating LLMs in social science: Epistemic threats and emerging norms5harxivHow Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism5h
The Bubble Brief
WEEKLY

Read software engineering insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews