·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing4h◆Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5h◆Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex5h◆Market Design for AI: Beyond the Copyright Binary5h◆Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents5h◆TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-25h◆DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning5h◆From World Models to World Action Models: A Concise Tutorial for Robotics5h◆QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting5h◆Multi-Turn On-Policy Distillation with Prefix Replay5h◆Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary5h◆Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts5h◆MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents5h◆PhantomFill: When the Form Demands an Answer, Language Models Invent One5h◆Error Certificates for KV-Cache Eviction via Randomized Design5h◆Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution5h◆MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities5h◆LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction5h◆Mwando: Leveraging AI to Preserve and Teach shiKomori5h◆The JEPA Paradox in Language: The Geometry of Linguistic Alternatives5h◆Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing4h◆Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5h◆Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex5h◆Market Design for AI: Beyond the Copyright Binary5h◆Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents5h◆TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-25h◆DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning5h◆From World Models to World Action Models: A Concise Tutorial for Robotics5h◆QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting5h◆Multi-Turn On-Policy Distillation with Prefix Replay5h◆Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary5h◆Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts5h◆MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents5h◆PhantomFill: When the Form Demands an Answer, Language Models Invent One5h◆Error Certificates for KV-Cache Eviction via Randomized Design5h◆Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution5h◆MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities5h◆LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction5h◆Mwando: Leveraging AI to Preserve and Teach shiKomori5h◆The JEPA Paradox in Language: The Geometry of Linguistic Alternatives5h◆
Tag

#synthetic-data

3 articles tagged #synthetic-data

arxivJul 10

Workload-Preserving Differentially Private Synthetic Data for Causal Inference via Maximum-Entropy Calibration

arXiv:2607.08122v1 Announce Type: new Abstract: Workload-based differentially private (DP) synthetic data methods privately measure aggregate queries and post-process the noisy answers into synthetic records. Generic workloads can achieve strong distributional fidelity, but causal estimands such as

#synthetic-data#differential-privacy#causal-inferenceRead on arxiv →
arxivMay 29

HomeModelsNews
Error as a Lens: Probing LLM Reasoning through Synthetic Misconception Generation

arXiv:2605.29007v1 Announce Type: new Abstract: Personalized tutoring, teacher training, and education research need access to \emph{targeted} synthetic misconceptions, but privacy and IRB constraints make labelled corpora of real student errors scarce. LLMs could in principle generate synthetic err

LL1 model#education#synthetic-data#language-modelsRead on arxiv →
arxivApr 29bullish

STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator

arXiv:2604.24544v1 Announce Type: new Abstract: The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific evaluation datasets; however, the collection of such datasets is challenging due to privacy concerns, re

LATG2 models#benchmark#evaluation#language-modelsRead on arxiv →