collaborators

8 papers

cs.LG2025

A Variational Manifold Embedding Framework for Nonlinear Dimensionality Reduction

John J. Vastola, Samuel J. Gershman, Kanaka Rajan

Dimensionality reduction algorithms like principal component analysis (PCA) are workhorses of machine learning and neuroscience, but each has well-known limitations. Variants of PC…

cs.LG2025

Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules

John J. Vastola, Samuel J. Gershman, Kanaka Rajan

Learning rules -- prescriptions for updating model parameters to improve performance -- are typically assumed rather than derived. Why do some learning rules work better than other…

cs.LG2025

A circuit for predicting hierarchical structure in-context in Large Language Models

Tankred Saanum, Can Demircan, Samuel J. Gershman +1

Large Language Models (LLMs) excel at in-context learning, the ability to use information provided as context to improve prediction of future tokens. Induction heads have been argu…

cs.AI2025

Assessing Adaptive World Models in Machines with Novel Games

Lance Ying, Katherine M. Collins, Prafull Sharma +11

Human intelligence exhibits a remarkable capacity for rapid adaptation and effective problem-solving in novel and unfamiliar contexts. We argue that this profound adaptability is f…

cs.LG2025

Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers

Kazuki Irie, Morris Yau, Samuel J. Gershman

We develop hybrid memory architectures for general-purpose sequence processing neural networks, that combine key-value memory using softmax attention (KV-memory) with fast weight m…

q-bio.NC2025

Key-value memory in the brain

Samuel J. Gershman, Ila Fiete, Kazuki Irie

Classical models of memory in psychology and neuroscience rely on similarity-based retrieval of stored patterns, where similarity is a function of retrieval cues and the stored pat…