arxiv
PublishedJuly 18, 2026 at 4:00 AM
▲bullish
RAD: Retrieval High-quality Demonstrations to Enhance Decision-making
Publisher summary· verbatim
arXiv:2507.15356v2 Announce Type: replace Abstract: Offline reinforcement learning (RL) learns policies from fixed datasets, thereby avoiding costly or unsafe environment interactions. However, its reliance on finite static datasets inherently restricts the ability to generalize beyond the training
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9harxivSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks9harxivDistribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts9harxivIn RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning9hThe Bubble Brief
WEEKLYRead reinforcement-learning insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗