·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
UN says AI safeguards can’t wait for certainty2h◆Amazon doesn’t trust Meta’s Muse AI agent3h◆LoRA Enhanced Contrastive Learning with SAS Vision Transformers8h◆Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency8h◆TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar8h◆Runtime Authorization for Resources Acquired by AI Agents8h◆JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems8h◆Multi-Domain Clustering via Measure Quantization8h◆Generative inversion for early ranking of competing geologic interpretations8h◆Data-free On-policy Distillation8h◆Dynamic Lagging using Stable-Prefix Training for Simultaneous Translation8h◆Do small language models know what they don't know?8h◆CaLR: Causal Latent Revision for Robust Diffusion Reasoning8h◆Can Agents Design Better Chips with a Higher Level Abstraction?8h◆SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity8h◆AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture8h◆Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation8h◆One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction8h◆GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation8h◆Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation8h◆UN says AI safeguards can’t wait for certainty2h◆Amazon doesn’t trust Meta’s Muse AI agent3h◆LoRA Enhanced Contrastive Learning with SAS Vision Transformers8h◆Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency8h◆TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar8h◆Runtime Authorization for Resources Acquired by AI Agents8h◆JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems8h◆Multi-Domain Clustering via Measure Quantization8h◆Generative inversion for early ranking of competing geologic interpretations8h◆Data-free On-policy Distillation8h◆Dynamic Lagging using Stable-Prefix Training for Simultaneous Translation8h◆Do small language models know what they don't know?8h◆CaLR: Causal Latent Revision for Robust Diffusion Reasoning8h◆Can Agents Design Better Chips with a Higher Level Abstraction?8h◆SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity8h◆AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture8h◆Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation8h◆One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction8h◆GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation8h◆Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation8h◆
News/Data-free On-policy Distillation
arxiv
PublishedSeptember 21, 2026 at 4:00 AM
—neutral

Data-free On-policy Distillation

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2609.14193v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined. On the two teacher--student pairings most common in practice, w

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivLoRA Enhanced Contrastive Learning with SAS Vision Transformers8harxivHallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency8harxivTatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar8harxivRuntime Authorization for Resources Acquired by AI Agents8h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews