Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Building Better Activation Oracles
Jan Bauer, Celeste De Schamphelaere, Adam Karvonen +2
Activation Oracles (AOs) are promising methods for interpreting residual stream activations. However, current AOs face important issues, such as hallucinations and vagueness. Addit…
cs.LG2026
A unified theory of feature learning in RNNs and DNNs
Jan P. Bauer, Kirsten Fischer, Moritz Helias +1
Recurrent and deep neural networks (RNNs/DNNs) are cornerstone architectures in machine learning. Remarkably, RNNs differ from DNNs only by weight sharing, as can be shown through…