Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Samuel Soo, Chen Guang, Wesley Teng +3
Effective and reliable control over large language model (LLM) behavior is a significant challenge. While activation steering methods, which add steering vectors to a model's hidde…
cs.LG2020
Causal Explanations of Image Misclassifications
Yan Min, Miles Bennett
The causal explanation of image misclassifications is an understudied niche, which can potentially provide valuable insights in model interpretability and increase prediction accur…