3 papers
cs.LG2025
Distribution-Aware Feature Selection for SAEs
Narmeen Oozeer, Nirmalendu Prakash, Michael Lan +2
Sparse autoencoders (SAEs) decompose neural activations into interpretable features. A widely adopted variant, the TopK SAE, reconstructs each token from its K most active latents.…
cs.LG2024
Bilinear Convolution Decomposition for Causal RL Interpretability
Narmeen Oozeer, Sinem Erisken, Alice Rigg
Efforts to interpret reinforcement learning (RL) models often rely on high-level techniques such as attribution or probing, which provide only correlational insights and coarse cau…
cs.LG2024
Weight-based Decomposition: A Case for Bilinear MLPs
Michael T. Pearce, Thomas Dooms, Alice Rigg
Gated Linear Units (GLUs) have become a common building block in modern foundation models. Bilinear layers drop the non-linearity in the "gate" but still have comparable performanc…