8 papers
A Variational Manifold Embedding Framework for Nonlinear Dimensionality Reduction
John J. Vastola, Samuel J. Gershman, Kanaka Rajan
Dimensionality reduction algorithms like principal component analysis (PCA) are workhorses of machine learning and neuroscience, but each has well-known limitations. Variants of PC…
Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
John J. Vastola, Samuel J. Gershman, Kanaka Rajan
Learning rules -- prescriptions for updating model parameters to improve performance -- are typically assumed rather than derived. Why do some learning rules work better than other…
A circuit for predicting hierarchical structure in-context in Large Language Models
Tankred Saanum, Can Demircan, Samuel J. Gershman +1
Large Language Models (LLMs) excel at in-context learning, the ability to use information provided as context to improve prediction of future tokens. Induction heads have been argu…
Assessing Adaptive World Models in Machines with Novel Games
Lance Ying, Katherine M. Collins, Prafull Sharma +11
Human intelligence exhibits a remarkable capacity for rapid adaptation and effective problem-solving in novel and unfamiliar contexts. We argue that this profound adaptability is f…
Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
Kazuki Irie, Morris Yau, Samuel J. Gershman
We develop hybrid memory architectures for general-purpose sequence processing neural networks, that combine key-value memory using softmax attention (KV-memory) with fast weight m…
Key-value memory in the brain
Samuel J. Gershman, Ila Fiete, Kazuki Irie
Classical models of memory in psychology and neuroscience rely on similarity-based retrieval of stored patterns, where similarity is a function of retrieval cues and the stored pat…