activity
20182022
most citedLinguistically-Informed Transformations (LIT): A Method for Automatically Generating Contrast Sets

4 citations · 6 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2022

Testing Pre-trained Language Models' Understanding of Distributivity via Causal Mediation Analysis

Pangbo Ban, Yifan Jiang, Tianran Liu +1

To what extent do pre-trained language models grasp semantic knowledge regarding the phenomenon of distributivity? In this paper, we introduce DistNLI, a new diagnostic dataset for…

cs.CL2022

Probing for Understanding of English Verb Classes and Alternations in Large Pre-trained Language Models

David K. Yi, James V. Bruno, Jiayu Han +2

We investigate the extent to which verb alternation classes, as described by Levin (1993), are encoded in the embeddings of Large Pre-trained Language Models (PLMs) such as BERT, R…

cs.CL2021

Language Models Use Monotonicity to Assess NPI Licensing

Jaap Jumelet, Milica Denić, Jakub Szymanik +2

We investigate the semantic knowledge of language models (LMs), focusing on (1) whether these LMs create categories of linguistic environments based on their semantic monotonicity…

cs.CL2021

A multilabel approach to morphosyntactic probing

Naomi Tachikawa Shapiro, Amandalynne Paullada, Shane Steinert-Threlkeld

We introduce a multilabel probing task to assess the morphosyntactic representations of word embeddings from multilingual language models. We demonstrate this task with multilingua…

cs.CL2021

A Masked Segmental Language Model for Unsupervised Natural Language Segmentation

C. M. Downey, Fei Xia, Gina-Anne Levow +1

Segmentation remains an important preprocessing step both in languages where "words" or other important syntactic/semantic units (like morphemes) are not clearly delineated by whit…

cs.CL20204 cited

Linguistically-Informed Transformations (LIT): A Method for Automatically Generating Contrast Sets

Chuanrong Li, Lin Shengshuo, Leo Z. Liu +3

Although large-scale pretrained language models, such as BERT and RoBERTa, have achieved superhuman performance on in-distribution test sets, their performance suffers on out-of-di…