arxiv
PublishedJuly 15, 2026 at 4:00 AM
—neutral
Toward Localizing and Repairing Bias in Transformer Attention Heads
Publisher summary· verbatim
arXiv:2607.12863v1 Announce Type: cross Abstract: Transformer language models are increasingly used as software components, yet biased outputs remain difficult to localize and repair inside the model. Existing fairness testing and repair methods largely operate at the input-output or retraining leve
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning4harxivSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks4harxivDistribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts4harxivIn RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning4hThe Bubble Brief
WEEKLYRead fairness insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗