arxiv
PublishedSeptember 11, 2026 at 4:00 AM
MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions
Publisher summary· verbatim
arXiv:2609.11253v1 Announce Type: cross Abstract: Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpredictable,
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning8harxivSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks8harxivDistribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts8harxivIn RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning8hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗