arxiv
PublishedSeptember 7, 2026 at 4:00 AM
—neutral
SharedSAE: One Feature Dictionary Across Language Models
Publisher summary· verbatim
arXiv:2609.04344v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here, we show that a single shared SAE can replace a collection of dedicated per-model S
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivExtremely Sparse Supervision Incentivizes Reasoning Ability1darxivConstructing and Evaluating Clinical Reasoning Trajectories for Medical Agent1darxivWhen Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs1darxivBlockchain-Enabled Secure Logging for Fiscal Electronic Mechanisms: Evaluation of the Greek eSEND and myDATA Tax Systems1dThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗