16 citations · 16 across the 1 of their papers we have counts for
2 papers
cs.CL2020★ 16 cited
When Do You Need Billions of Words of Pretraining Data?
Yian Zhang, Alex Warstadt, Haau-Sing Li +1
NLP is currently dominated by general-purpose pretrained language models like RoBERTa, which achieve strong performance on NLU tasks through pretraining on billions of words. But w…
cs.CL2020
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)
Alex Warstadt, Yian Zhang, Haau-Sing Li +2
One reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding. However, we want pretrained…