activity
20182021
most citedFine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping

215 citations · 350 across the 5 of their papers we have counts for

collaborators

16 papers

cs.CL2021

Provable Limitations of Acquiring Meaning from Ungrounded Form: What Will Future Language Models Understand?

William Merrill, Yoav Goldberg, Roy Schwartz +1

Language models trained on billions of tokens have recently led to unprecedented results on many NLP tasks. This success raises the question of whether, in principle, a system can…

cs.CL2021121 cited

Random Feature Attention

Hao Peng, Nikolaos Pappas, Dani Yogatama +3

Transformers are state-of-the-art models for a variety of sequence modeling tasks. At their core is an attention function which models pairwise interactions between the inputs at e…

cs.CL2021

Automatic Generation of Contrast Sets from Scene Graphs: Probing the Compositional Consistency of GQA

Yonatan Bitton, Gabriel Stanovsky, Roy Schwartz +1

Recent works have shown that supervised models often exploit data artifacts to achieve good test scores while their performance severely degrades on samples outside their training…

cs.CL2020

Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics

Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie +4

Large datasets have become commonplace in NLP research. However, the increased emphasis on data quantity has made it challenging to assess the quality of data. We introduce Data Ma…

cs.CL2020

Extracting a Knowledge Base of Mechanisms from COVID-19 Papers

Tom Hope, Aida Amini, David Wadden +6

The COVID-19 pandemic has spawned a diverse body of scientific literature that is challenging to navigate, stimulating interest in automated tools to help find useful knowledge. We…

cs.CL20202 cited

A Mixture of Heads is Better than Heads

Hao Peng, Roy Schwartz, Dianqi Li +1

Multi-head attentive neural architectures have achieved state-of-the-art results on a variety of natural language processing tasks. Evidence has shown that they are overparameteriz…