4 papers
Gradient Descent as Implicit EM in Distance-Based Neural Models
Alan Oursland
Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tr…
Deriving Decoder-Free Sparse Autoencoders from First Principles
Alan Oursland
Gradient descent on log-sum-exp (LSE) objectives performs implicit expectation--maximization (EM): the gradient with respect to each component output equals its responsibility. The…
Neural Networks Learn Distance Metrics
Alan Oursland
Neural networks may naturally favor distance-based representations, where smaller activations indicate closer proximity to learned prototypes. This contrasts with intensity-based a…
Neural Networks Use Distance Metrics
Alan Oursland
We present empirical evidence that neural networks with ReLU and Absolute Value activations learn distance-based representations. We independently manipulate both distance and inte…