3 papers
cs.LG2026
Deriving Decoder-Free Sparse Autoencoders from First Principles
Alan Oursland
Gradient descent on log-sum-exp (LSE) objectives performs implicit expectation--maximization (EM): the gradient with respect to each component output equals its responsibility. The…
cs.LG2025
Neural Networks Learn Distance Metrics
Alan Oursland
Neural networks may naturally favor distance-based representations, where smaller activations indicate closer proximity to learned prototypes. This contrasts with intensity-based a…
cs.LG2024
Neural Networks Use Distance Metrics
Alan Oursland
We present empirical evidence that neural networks with ReLU and Absolute Value activations learn distance-based representations. We independently manipulate both distance and inte…