139 citations · 376 across the 19 of their papers we have counts for
6 papers · 2 filters
Adversarially Constructed Evaluation Sets Are More Challenging, but May Not Be Fair
Jason Phang, Angelica Chen, William Huang +1
More capable language models increasingly saturate existing task benchmarks, in some cases outperforming humans. This has left little headroom with which to measure further progres…
Clean or Annotate: How to Spend a Limited Data Collection Budget
Derek Chen, Zhou Yu, Samuel R. Bowman
Crowdsourcing platforms are often used to collect datasets for training machine learning models, despite higher levels of inaccurate labeling compared to expert labeling. There are…
Fine-Tuned Transformers Show Clusters of Similar Representations Across Layers
Jason Phang, Haokun Liu, Samuel R. Bowman
Despite the success of fine-tuning pretrained language encoders like BERT for downstream natural language understanding (NLU) tasks, it is still poorly understood how neural networ…
Does Putting a Linguist in the Loop Improve NLU Data Collection?
Alicia Parrish, William Huang, Omar Agha +7
Many crowdsourced NLP datasets contain systematic gaps and biases that are identified only after data collection is complete. Identifying these issues from early data samples durin…
Efficient transfer learning for NLP with ELECTRA
François Mercier
Clark et al. [2020] claims that the ELECTRA approach is highly efficient in NLP performances relative to computation budget. As such, this reproducibility study focus on this claim…
What Will it Take to Fix Benchmarking in Natural Language Understanding?
Samuel R. Bowman, George E. Dahl
Evaluation for many natural language understanding (NLU) tasks is broken: Unreliable and biased systems score so highly on standard benchmarks that there is little room for researc…