565 citations · 730 across the 26 of their papers we have counts for
4 papers · 2 filters
When Do You Need Billions of Words of Pretraining Data?
Yian Zhang, Alex Warstadt, Haau-Sing Li +1
NLP is currently dominated by general-purpose pretrained language models like RoBERTa, which achieve strong performance on NLU tasks through pretraining on billions of words. But w…
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)
Alex Warstadt, Yian Zhang, Haau-Sing Li +2
One reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding. However, we want pretrained…
Can neural networks acquire a structural bias from raw linguistic data?
Alex Warstadt, Samuel R. Bowman
We evaluate whether BERT, a widely used neural network for sentence processing, acquires an inductive bias towards forming structural generalizations through pretraining on raw dat…
Are Natural Language Inference Models IMPPRESsive? Learning IMPlicature and PRESupposition
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan +1
Natural language inference (NLI) is an increasingly important task for natural language understanding, which requires one to infer whether a sentence entails another. However, the…