arxiv
PublishedSeptember 4, 2026 at 4:00 AM
Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks
Publisher summary· verbatim
arXiv:2608.25390v2 Announce Type: replace-cross Abstract: Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivCulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning14harxivProactive Service Agents: A Unified Decision Framework, Methods, and Evaluation14harxivX-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System14harxivA Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors14hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗