2 papers
cs.LG2026
An evolutionary perspective on modes of learning in Transformers
Alexander Y. Ku, Thomas L. Griffiths, Stephanie C. Y. Chan
The success of Transformers lies in their ability to improve inference through two complementary strategies: the permanent refinement of model parameters via in-weight learning (IW…
cs.CL2025
On the generalization of language models from in-context learning and finetuning: a controlled study
Andrew K. Lampinen, Arslan Chaudhry, Stephanie C. Y. Chan +7
Large language models exhibit exciting capabilities, yet can show surprisingly narrow generalization from finetuning. E.g. they can fail to generalize to simple reversals of relati…