51 citations · 51 across the 3 of their papers we have counts for
3 papers
cs.LG2023★ 51 cited
Sparse Autoencoders Find Highly Interpretable Features in Language Models
Hoagy Cunningham, Aidan Ewart, Logan Riggs +2
One of the roadblocks to a better understanding of neural networks' internals is \textit{polysemanticity}, where neurons appear to activate in multiple, semantically distinct conte…
cs.LG2023
Attention-Only Transformers and Implementing MLPs with Attention Heads
Robert Huben, Valerie Morris
The transformer architecture is widely used in machine learning models and consists of two alternating sublayers: attention heads and MLPs. We prove that an MLP neuron can be imple…
math.GR2021
Gauge-Invariant Uniqueness and Reductions of Ordered Groups
Robert Huben
A reduction of an ordered group to another ordered group is an order homomorphism which maps each interval bijectively onto . We show that if …