arxiv
PublishedSeptember 1, 2026 at 4:00 AM
Do VLMs Share Safety Neurons Across Modalities?
Publisher summary· verbatim
arXiv:2608.30750v1 Announce Type: new Abstract: Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, ho
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
The Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗