8 papers
Predicting the Emergence of Induction Heads in Language Model Pretraining
Tatsuya Aoyama, Ethan Gotlieb Wilcox, Nathan Schneider
Specialized attention heads dubbed induction heads (IHs) have been argued to underlie the remarkable in-context learning capabilities of modern language models; yet, a precise char…
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs
Xiulin Yang, Tatsuya Aoyama, Yuekun Yao +1
Do language models (LMs) offer insights into human language learning? A common argument against this idea is that because their architecture and training paradigm are so vastly dif…
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
Wesley Scivetti, Ethan Wilcox, Nathan Schneider +2
Grasping the semantics of rare constructions (form-meaning pairings) has been shown to be a challenging problem that has currently only been solved by the largest LLMs. It remains…
Information-Theoretic Storage Cost in Sentence Comprehension
Kohei Kajikawa, Shinnosuke Isono, Ethan Gotlieb Wilcox
Real-time sentence comprehension imposes a significant load on working memory, as comprehenders must maintain contextual information to anticipate future input. While measures of s…
Function Words as Statistical Cues for Language Learning
Xiulin Yang, Heidi Getz, Ethan Gotlieb Wilcox
What statistical properties might support learning abstract grammatical knowledge from linear input? We address this question by examining the statistical distribution of function…
A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models
Xiulin Yang, Arianna Bisazza, Nathan Schneider +1
Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using Artificial Neural Networks (ANNs). The results suggest that ANN-based language model…