12 citations · 22 across the 8 of their papers we have counts for
4 papers · 1 filter
Principles from Clinical Research for NLP Model Generalization
Aparna Elangovan, Jiayuan He, Yuan Li +1
The NLP community typically relies on performance of a model on a held-out test set to assess generalization. Performance drops observed in datasets outside of official test sets a…
Effects of Human Adversarial and Affable Samples on BERT Generalization
Aparna Elangovan, Jiayuan He, Yuan Li +1
BERT-based models have had strong performance on leaderboards, yet have been demonstrably worse in real-world settings requiring generalization. Limited quantities of training data…
Collective Human Opinions in Semantic Textual Similarity
Yuxia Wang, Shimin Tao, Ning Xie +3
Despite the subjective nature of semantic textual similarity (STS) and pervasive disagreements in STS annotation, existing benchmarks have used averaged human ratings as the gold s…
Language models are not naysayers: An analysis of language models on negation benchmarks
Thinh Hung Truong, Timothy Baldwin, Karin Verspoor +1
Negation has been shown to be a major bottleneck for masked language models, such as BERT. However, whether this finding still holds for larger-sized auto-regressive language model…