·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
OpenAI, Anthropic, Google have been in talks on AI safety for weeks43m◆AEO startup Profound hits unicorn valuation, raises $180M Series D 7 months after last round1h◆Meta’s new One subscriptions put a price on social media and AI1h◆Former TikTok execs built an app that uses AI to teach you how to pose for a photo1h◆Discover how to take your startup from prototype to production at TechCrunch Disrupt 20262h◆4 days left to exhibit at TechCrunch Disrupt 20262h◆This doorbell camera lets a human security guard watch your front door2h◆Early Anthropic hire, former METR COO have found a way to rein in rogue AI agents3h◆New insights from Google’s AI & Economy ATLAS3h◆Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear4h◆What’s at stake in AI’s trillion-dollar gamble6h◆JaxAHT: A JAX-Based Library for Ad Hoc Teamwork12h◆MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents12h◆Signatures of Steerability in Activation Space of Language Models12h◆Data-free On-policy Distillation12h◆A latent dimension of Condorcet's jury theorem for multiple AI advisers12h◆Personalizing Personal Health Interfaces: Co-Design with Generative AI12h◆Ensemble Complexity in Photovoltaic Forecasting12h◆Self-Evolving Memory for Generative Recommendation12h◆Mecha-nudges for Machines12h◆OpenAI, Anthropic, Google have been in talks on AI safety for weeks43m◆AEO startup Profound hits unicorn valuation, raises $180M Series D 7 months after last round1h◆Meta’s new One subscriptions put a price on social media and AI1h◆Former TikTok execs built an app that uses AI to teach you how to pose for a photo1h◆Discover how to take your startup from prototype to production at TechCrunch Disrupt 20262h◆4 days left to exhibit at TechCrunch Disrupt 20262h◆This doorbell camera lets a human security guard watch your front door2h◆Early Anthropic hire, former METR COO have found a way to rein in rogue AI agents3h◆New insights from Google’s AI & Economy ATLAS3h◆Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear4h◆What’s at stake in AI’s trillion-dollar gamble6h◆JaxAHT: A JAX-Based Library for Ad Hoc Teamwork12h◆MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents12h◆Signatures of Steerability in Activation Space of Language Models12h◆Data-free On-policy Distillation12h◆A latent dimension of Condorcet's jury theorem for multiple AI advisers12h◆Personalizing Personal Health Interfaces: Co-Design with Generative AI12h◆Ensemble Complexity in Photovoltaic Forecasting12h◆Self-Evolving Memory for Generative Recommendation12h◆Mecha-nudges for Machines12h◆
News/Data-free On-policy Distillation
arxiv
PublishedSeptember 15, 2026 at 4:00 AM
—neutral

Data-free On-policy Distillation

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2609.14193v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined. On the two teacher-student pairings most common in practice, we find OP

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivJaxAHT: A JAX-Based Library for Ad Hoc Teamwork12harxivMOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents12harxivSignatures of Steerability in Activation Space of Language Models12harxivA latent dimension of Condorcet's jury theorem for multiple AI advisers12h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews