arxiv
PublishedJuly 28, 2026 at 4:00 AM
Reference Feature Atlases for Mechanistic Auditing of Language Models
Publisher summary· verbatim
arXiv:2607.22570v1 Announce Type: new Abstract: Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by f
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivAgentic Permissions Policy Algebra for Taint Confinement in LLM Agents2harxivSparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects2harxivBeyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks2harxivThe One-Word Census: Answer-Choice Conformity Across 44 Language Models2hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗