arxiv
PublishedSeptember 17, 2026 at 4:00 AM
No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback
Publisher summary· verbatim
arXiv:2609.17550v1 Announce Type: new Abstract: Language models frequently abandon correct answers when users push back. We study this in two small instruction-tuned models from different families, Qwen2.5-1.5B and Llama-3.2-1B, over TriviaQA: the model answers, is challenged with one of four script
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivGeospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines36marxivVisual Cue Guided Video Planning for Generalizable Robot Navigation36marxivA unified framework for global and local interpretability using adaptive derivative-ordered random explanation36marxivAfter the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem36mThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗