9 papers
Bilinear autoencoders find interpretable manifolds
Thomas Dooms, Ward Gauderis, Geraint Wiggins +1
Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span manifolds that current linea…
TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement
Arian Sabaghi, José Oramas
Weakly supervised object localization (WSOL) aims to localize target objects in images using only image-level labels. Despite recent progress, many approaches still rely on multi-s…
Explainability-Driven Dimensionality Reduction for Hyperspectral Imaging
Salma Haidar, José Oramas
Hyperspectral imaging (HSI) provides rich spectral information for precise material classification and analysis; however, its high dimensionality introduces a computational burden…
Bilinear MLPs enable weight-based mechanistic interpretability
Michael T. Pearce, Thomas Dooms, Alice Rigg +2
A mechanistic understanding of how MLPs do computation in deep neural networks remains elusive. Current interpretability work can extract features from hidden activations over an i…
Smooth InfoMax -- Towards Easier Post-Hoc Interpretability
Fabian Denoodt, Bart de Boer, José Oramas
We introduce Smooth InfoMax (SIM), a self-supervised representation learning method that incorporates interpretability constraints into the latent representations at different dept…
Compositionality Unlocks Deep Interpretable Models
Thomas Dooms, Ward Gauderis, Geraint A. Wiggins +1
We propose -net, an intrinsically interpretable architecture combining the compositional multilinear structure of tensor networks with the expressivity and efficiency of deep n…