2 papers
cs.CL2025
On the generalization of language models from in-context learning and finetuning: a controlled study
Andrew K. Lampinen, Arslan Chaudhry, Stephanie C. Y. Chan +7
Large language models exhibit exciting capabilities, yet can show surprisingly narrow generalization from finetuning. E.g. they can fail to generalize to simple reversals of relati…
cs.LG2024
Uncovering Layer-Dependent Activation Sparsity Patterns in ReLU Transformers
Cody Wild, Jesper Anderson
Previous work has demonstrated that MLPs within ReLU Transformers exhibit high levels of sparsity, with many of their activations equal to zero for any given token. We build on tha…