arxiv
PublishedSeptember 10, 2026 at 4:00 AM
Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
Publisher summary· verbatim
arXiv:2609.07627v1 Announce Type: new Abstract: AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivPlanning and Scheduling Business Processes under Control-Flow Uncertainty15harxivSIM: Subspace Interaction-based Method for Token-Level Text Anomaly Detection15harxivWAPP: Safe Learning of Positive Security WAF Policies from Live Traffic15harxivPAN: A World Model for General, Actionable, and Long-Horizon World Simulation15hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗