11 papers
Linguistic Context Recodes Visual Representations in Vision-Language Models
Brian Song, Michael A. Lepori, Ellie Pavlick
Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though visi…
Transferring Linear Features Across Language Models With Model Stitching
Alan Chen, Jack Merullo, Alessandro Stolfo +1
In this work, we demonstrate that affine mappings between residual streams of language models is a cheap way to effectively transfer represented features between models. We apply t…
$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources
Apoorv Khandelwal, Tian Yun, Nihal V. Nayak +4
Pre-training is notoriously compute-intensive and academic researchers are notoriously under-resourced. It is, therefore, commonly assumed that academics can't pre-train models. In…
Deep Neural Networks Can Learn Generalizable Same-Different Visual Relations
Alexa R. Tartaglini, Sheridan Feucht, Michael A. Lepori +4
Although deep neural networks can achieve human-level performance on many object recognition benchmarks, prior work suggests that these same models fail to learn simple abstract re…
Whither symbols in the era of advanced neural networks?
Thomas L. Griffiths, Brenden M. Lake, R. Thomas McCoy +2
Some of the strongest evidence that human minds should be thought about in terms of symbolic systems has been the way they combine ideas, produce novelty, and learn quickly. We arg…
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
Although it is known that transformer language models (LMs) pass features from early layers to later layers, it is not well understood how this information is represented and route…