5 papers
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
Nicolas Anguita, Francesco Locatello, Andrew M. Saxe +4
Pretraining and fine-tuning are central stages in modern machine learning systems. In practice, feature learning plays an important role across both stages: deep neural networks le…
A mathematical theory of balancing relational generalization and memorization
Luke Cheng, Samuel Lippl
Humans, animals, and modern machine learning models exhibit impressive abilities to learn complex behaviors and generalize these behaviors to unseen situations. This ability requir…
Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
Samuel Lippl, Thomas McGee, Kimberly Lopez +5
How do latent and inference time computations enable large language models (LLMs) to solve multi-step reasoning? We introduce a framework for tracing and steering algorithmic primi…
When does compositional structure yield compositional generalization? A kernel theory
Samuel Lippl, Kim Stachenfeld
Compositional generalization (the ability to respond correctly to novel combinations of familiar components) is thought to be a cornerstone of intelligent behavior. Compositionally…
Inductive biases of multi-task learning and finetuning: multiple regimes of feature reuse
Samuel Lippl, Jack W. Lindsey
Neural networks are often trained on multiple tasks, either simultaneously (multi-task learning, MTL) or sequentially (pretraining and subsequent finetuning, PT+FT). In particular,…