·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance2h◆Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study2h◆Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents2h◆Making Duplicate Reimbursement Unrepresentable: A Verified Ethereum E-Invoice System for Humans and AI Agents2h◆Is manual software optimization a thing of the past?2h◆HandAnthro: Automated Hand Anthropometry from a Single Image2h◆It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them2h◆AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems2h◆Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR2h◆ExceptionDrive: A Planning-Oriented Counterfactual Corner-Case Benchmark for Autonomous Driving2h◆Boids of a Feather Flock Together - Evolving Prey Behaviours Under Different Predator Attack Strategies2h◆Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction2h◆Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v32h◆The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment2h◆RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust2h◆Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation2h◆Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark2h◆Dagger: Decoupling-based Model Stealing Attack against Graph Neural Networks2h◆On Trajectory-Aware Training for Masked Diffusion Language Models2h◆$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient2h◆Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance2h◆Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study2h◆Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents2h◆Making Duplicate Reimbursement Unrepresentable: A Verified Ethereum E-Invoice System for Humans and AI Agents2h◆Is manual software optimization a thing of the past?2h◆HandAnthro: Automated Hand Anthropometry from a Single Image2h◆It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them2h◆AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems2h◆Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR2h◆ExceptionDrive: A Planning-Oriented Counterfactual Corner-Case Benchmark for Autonomous Driving2h◆Boids of a Feather Flock Together - Evolving Prey Behaviours Under Different Predator Attack Strategies2h◆Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction2h◆Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v32h◆The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment2h◆RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust2h◆Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation2h◆Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark2h◆Dagger: Decoupling-based Model Stealing Attack against Graph Neural Networks2h◆On Trajectory-Aware Training for Masked Diffusion Language Models2h◆$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient2h◆
News/Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
arxiv
PublishedSeptember 30, 2026 at 4:00 AM

Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2609.37868v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts i

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivPredictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance2harxivPixel-Level Transformers in Remote Sensing: A Canopy Height Case Study2harxivExplore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents2harxivBoosting Adversarial Robustness and Generalization with Dictionary Structure2h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
Built by Marouane Gazouzi
HomeModelsNews