16 citations · 36 across the 9 of their papers we have counts for
10 papers · 1 filter
Coloring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models
Aaron Mueller, Robert Frank, Tal Linzen +2
Relations between words are governed by hierarchical structure rather than linear ordering. Sequence-to-sequence (seq2seq) models, despite their success in downstream NLP applicati…
Do Language Models Learn Position-Role Mappings?
Jackson Petty, Michael Wilson, Robert Frank
How is knowledge of position-role mappings in natural language learned? We explore this question in a computational setting, testing whether a variety of well-performing pertained…
Transformers Generalize Linearly
Jackson Petty, Robert Frank
Natural language exhibits patterns of hierarchically governed dependencies, in which relations between words are sensitive to syntactic structure rather than linear ordering. While…
Sequence-to-Sequence Networks Learn the Meaning of Reflexive Anaphora
Robert Frank, Jackson Petty
Reflexive anaphora present a challenge for semantic interpretation: their meaning varies depending on context in a way that appears to require abstract variables. Past work has rai…
Does syntax need to grow on trees? Sources of hierarchical inductive bias in sequence-to-sequence networks
R. Thomas McCoy, Robert Frank, Tal Linzen
Learners that are exposed to the same training data might generalize differently due to differing inductive biases. In neural network models, inductive biases could in theory arise…
Detecting Syntactic Change Using a Neural Part-of-Speech Tagger
William Merrill, Gigi Felice Stark, Robert Frank
We train a diachronic long short-term memory (LSTM) part-of-speech tagger on a large corpus of American English from the 19th, 20th, and 21st centuries. We analyze the tagger's abi…