4 papers
From Attribution to Action: A Human-Centered Application of Activation Steering
Tobias Labarta, Maximilian Dreyer, Katharina Weitz +2
Explainable AI (XAI) methods reveal which features influence model predictions, yet provide limited means for practitioners to act on these explanations. Activation steering of com…
X-SYS: A Reference Architecture for Interactive Explanation Systems
Tobias Labarta, Nhi Hoang, Maximilian Dreyer +5
The explainable AI (XAI) research community has proposed numerous technical methods, yet deploying explainability as systems remains challenging: Interactive explanation systems re…
See What I Mean? CUE: A Cognitive Model of Understanding Explanations
Tobias Labarta, Nhi Hoang, Katharina Weitz +3
As machine learning systems increasingly inform critical decisions, the need for human-understandable explanations grows. Current evaluations of Explainable AI (XAI) often prioriti…
Mechanistic understanding and validation of large AI models with SemanticLens
Maximilian Dreyer, Jim Berend, Tobias Labarta +4
Unlike human-engineered systems such as aeroplanes, where each component's role and dependencies are well understood, the inner workings of AI models remain largely opaque, hinderi…