3 citations · 3 across the 9 of their papers we have counts for
Showing 2024 · cs.LGShow all
2 papers · 2 filters
cs.LG2024
Bilinear MLPs enable weight-based mechanistic interpretability
Michael T. Pearce, Thomas Dooms, Alice Rigg +2
A mechanistic understanding of how MLPs do computation in deep neural networks remains elusive. Current interpretability work can extract features from hidden activations over an i…
cs.LG2024
Weight-based Decomposition: A Case for Bilinear MLPs
Michael T. Pearce, Thomas Dooms, Alice Rigg
Gated Linear Units (GLUs) have become a common building block in modern foundation models. Bilinear layers drop the non-linearity in the "gate" but still have comparable performanc…