Publications (7)
Syntactic Surprisal From Neural Models Predicts, But Underestimates, Human Processing Difficulty From Syntactic Ambiguities
Suhas Arehalli, Brian Dillon, Tal Linzen
Humans exhibit garden path effects: When reading sentences that are temporarily structurally ambiguous, they slow down when the structure is disambiguated in favor of the less pref…
Interpretable dimensions support an effect of agentivity and telicity on split intransitivity
Eva Neu, Brian Dillon, Katrin Erk
Intransitive verbs fall into two different syntactic classes, unergatives and unaccusatives. It has long been argued that verbs describing an agentive action are more likely to app…
Human acceptability judgements for extractive sentence compression
Abram Handler, Brian Dillon, Brendan O'Connor
Recent approaches to English-language sentence compression rely on parallel corpora consisting of sentence-compression pairs. However, a sentence may be shortened in many different…
Simulating Human Memory with Language Models
Qihan Wang, Nicholas Tomlin, Michael Hu +2
Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic m…
Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis
William Timkey, Brian Dillon, Tal Linzen
Surprisal theory posits that the processing difficulty of a word is determined by its predictability in context, offering a potential link between human sentence processing and nex…
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
Farah Adeeba, Brian Dillon, Hassan Sajjad +1
Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages; however, they often include significantly less data for low-resource languages…
Evaluating Syntactic Properties of Seq2seq Output with a Broad Coverage HPSG: A Case Study on Machine Translation
Johnny Tian-Zheng Wei, Khiem Pham, Brian Dillon +1
Sequence to sequence (seq2seq) models are often employed in settings where the target output is natural language. However, the syntactic properties of the language generated from t…