activity
20232025
most citedWhat needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation

2 citations · 2 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2025

Softmax Linear: Transformers may learn to classify in-context by kernel gradient descent

Sara Dragutinović, Andrew M. Saxe, Aaditya K. Singh

The remarkable ability of transformers to learn new concepts solely by reading examples within the input prompt, termed in-context learning (ICL), is a crucial aspect of intelligen…

cs.LG2025

Distinct Computations Emerge From Compositional Curricula in In-Context Learning

Jin Hwa Lee, Andrew K. Lampinen, Aaditya K. Singh +1

In-context learning (ICL) research often considers learning a function in-context through a uniform sample of input-output pairs. Here, we investigate how presenting a compositiona…

cs.LG2025

Strategy Coopetition Explains the Emergence and Transience of In-Context Learning

Aaditya K. Singh, Ted Moskovitz, Sara Dragutinovic +3

In-context learning (ICL) is a powerful ability that emerges in transformer models, enabling them to learn from context without weight updates. Recent work has established emergent…

cs.LG2025

Nonlinear dynamics of localization in neural receptive fields

Leon Lufkin, Andrew M. Saxe, Erin Grant

Localized receptive fields -- neurons that are selective for certain contiguous spatiotemporal features of their input -- populate early sensory regions of the mammalian brain. Uns…

cs.LG20242 cited

What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation

Aaditya K. Singh, Ted Moskovitz, Felix Hill +2

In-context learning is a powerful emergent ability in transformer models. Prior work in mechanistic interpretability has identified a circuit element that may be critical for in-co…

cs.LG2023

The Transient Nature of Emergent In-Context Learning in Transformers

Aaditya K. Singh, Stephanie C. Y. Chan, Ted Moskovitz +3

Transformer neural networks can exhibit a surprising capacity for in-context learning (ICL) despite not being explicitly trained for it. Prior work has provided a deeper understand…