arxiv
PublishedApril 18, 2026 at 4:00 AM
—neutral
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?
Publisher summary· verbatim
arXiv:2509.12833v2 Announce Type: replace Abstract: Projection-based safety filters, which modify unsafe actions by mapping them to the closest safe alternative, are widely used to enforce safety constraints in reinforcement learning (RL). Two integration strategies are commonly considered: Safe env
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning3harxivSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks3harxivDistribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts3harxivIn RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning3hThe Bubble Brief
WEEKLYRead reinforcement-learning insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗