arxiv
PublishedSeptember 30, 2026 at 4:00 AM
—neutral
The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
Publisher summary· verbatim
arXiv:2609.37914v1 Announce Type: cross Abstract: Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations,
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivReasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions4harxivConsistent Plan-Act for Long-Horizon Agentic Tasks4harxivPredictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance4harxivBoosting Adversarial Robustness and Generalization with Dictionary Structure4hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗