arxiv
PublishedSeptember 29, 2026 at 4:00 AM
—neutral
Selective Off-Policy Reference Tuning with Plan Guidance
Publisher summary· verbatim
arXiv:2605.11505v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from t
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivReasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions11harxivConsistent Plan-Act for Long-Horizon Agentic Tasks11harxivPredictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance11harxivBoosting Adversarial Robustness and Generalization with Dictionary Structure11hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗