2 papers
cs.CL2026
Scaling Inherently Interpretable Language Models
Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail +7
Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult…
cs.CL2024
New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing
Andreas Madsen
As machine learning becomes more widespread and is used in more critical applications, it's important to provide explanations for these models, to prevent unintended behavior. Unfo…