arxiv
PublishedApril 16, 2026 at 4:00 AM
—neutral
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
Publisher summary· verbatim
arXiv:2509.25843v2 Announce Type: replace Abstract: Large language models (LLMs), despite being safety-aligned, exhibit brittle refusal behaviors that can be circumvented by simple linguistic changes. As tense jailbreaking demonstrates that models refusing harmful requests often comply when rephrase
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivAgentic Permissions Policy Algebra for Taint Confinement in LLM Agents2harxivSparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects2harxivBeyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks2harxivThe One-Word Census: Answer-Choice Conformity Across 44 Language Models2hThe Bubble Brief
WEEKLYRead safety insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗