19 citations · 50 across the 9 of their papers we have counts for
10 papers
Cost-effective Selection of Pretraining Data: A Case Study of Pretraining BERT on Social Media
Xiang Dai, Sarvnaz Karimi, Ben Hachey +1
Recent studies on domain-specific BERT models show that effectiveness on downstream tasks can be improved when models are pretrained on in-domain data. Often, the pretraining data…
Searching Scientific Literature for Answers on COVID-19 Questions
Vincent Nguyen, Maciek Rybinski, Sarvnaz Karimi +1
Finding answers related to a pandemic of a novel disease raises new challenges for information seeking and retrieval, as the new information becomes available gradually. TREC COVID…
An Effective Transition-based Model for Discontinuous NER
Xiang Dai, Sarvnaz Karimi, Ben Hachey +1
Unlike widely used Named Entity Recognition (NER) data sets in generic domains, biomedical NER data sets often contain mentions consisting of discontinuous spans. Conventional sequ…
Figurative Usage Detection of Symptom Words to Improve Personal Health Mention Detection
Adith Iyer, Aditya Joshi, Sarvnaz Karimi +2
Personal health mention detection deals with predicting whether or not a given sentence is a report of a health condition. Past work mentions errors in this prediction when symptom…
A Comparison of Word-based and Context-based Representations for Classification Problems in Health Informatics
Aditya Joshi, Sarvnaz Karimi, Ross Sparks +2
Distributed representations of text can be used as features when training a statistical classifier. These representations may be created as a composition of word vectors or as cont…
NNE: A Dataset for Nested Named Entity Recognition in English Newswire
Nicky Ringland, Xiang Dai, Ben Hachey +3
Named entity recognition (NER) is widely used in natural language processing applications and downstream tasks. However, most NER tools target flat annotation from popular datasets…