4 papers · 1 filter
SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable
Philip H. Lee, Parth Padalkar
Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post-hoc tools to decompose internal represen…
Symbolic Rule Extraction from Attention-Guided Sparse Representations in Vision Transformers
Parth Padalkar, Gopal Gupta
Recent neuro-symbolic approaches have successfully extracted symbolic rule-sets from CNN-based models to enhance interpretability. However, applying similar techniques to Vision Tr…
Improving Interpretability and Accuracy in Neuro-Symbolic Rule Extraction Using Class-Specific Sparse Filters
Parth Padalkar, Jaeseong Lee, Shiyi Wei +1
There has been significant focus on creating neuro-symbolic models for interpretable image classification using Convolutional Neural Networks (CNNs). These methods aim to replace t…
A Neurosymbolic Framework for Bias Correction in Convolutional Neural Networks
Parth Padalkar, Natalia Ålusarz, Ekaterina Komendantskaya +1
Recent efforts in interpreting Convolutional Neural Networks (CNNs) focus on translating the activation of CNN filters into a stratified Answer Set Program (ASP) rule-sets. The CNN…