18 citations · 34 across the 5 of their papers we have counts for
6 papers
Cost-effective Selection of Pretraining Data: A Case Study of Pretraining BERT on Social Media
Xiang Dai, Sarvnaz Karimi, Ben Hachey +1
Recent studies on domain-specific BERT models show that effectiveness on downstream tasks can be improved when models are pretrained on in-domain data. Often, the pretraining data…
An Effective Transition-based Model for Discontinuous NER
Xiang Dai, Sarvnaz Karimi, Ben Hachey +1
Unlike widely used Named Entity Recognition (NER) data sets in generic domains, biomedical NER data sets often contain mentions consisting of discontinuous spans. Conventional sequ…
NNE: A Dataset for Nested Named Entity Recognition in English Newswire
Nicky Ringland, Xiang Dai, Ben Hachey +3
Named entity recognition (NER) is widely used in natural language processing applications and downstream tasks. However, most NER tools target flat annotation from popular datasets…
Using Similarity Measures to Select Pretraining Data for NER
Xiang Dai, Sarvnaz Karimi, Ben Hachey +1
Word vectors and Language Models (LMs) pretrained on a large amount of unlabelled data can dramatically improve various Natural Language Processing (NLP) tasks. However, the measur…
Learning to generate one-sentence biographies from Wikidata
Andrew Chisholm, Will Radford, Ben Hachey
We investigate the generation of one-sentence Wikipedia biographies from facts derived from Wikidata slot-value pairs. We train a recurrent neural network sequence-to-sequence mode…
Post-edit Analysis of Collective Biography Generation
Bo Han, Will Radford, Anaïs Cadilhac +3
Text generation is increasingly common but often requires manual post-editing where high precision is critical to end users. However, manual editing is expensive so we want to ensu…