Publications (13)
Cross-Lingual Ability of Multilingual BERT: An Empirical Study
Karthikeyan K, Zihan Wang, Stephen Mayhew +1
Recent work has exhibited the surprising cross-lingual abilities of multilingual BERT (M-BERT) -- surprising since it is trained without any cross-lingual objective and with no ali…
Transliteration in Any Language with Surrogate Languages
Stephen Mayhew, Christos Christodoulopoulos, Dan Roth
We introduce a method for transliteration generation that can produce transliterations in every language. Where previous results are only as multilingual as Wikipedia, we show how…
Robust Named Entity Recognition with Truecasing Pretraining
Stephen Mayhew, Nitish Gupta, Dan Roth
Although modern named entity recognition (NER) systems show impressive performance on standard datasets, they perform poorly when presented with noisy data. In particular, capitali…
Named Entity Recognition with Partially Annotated Training Data
Stephen Mayhew, Snigdha Chaturvedi, Chen-Tse Tsai +1
Supervised machine learning assumes the availability of fully-labeled data, but in many cases, such as low-resource languages, the only data available is partially annotated. We st…
ner and pos when nothing is capitalized
Stephen Mayhew, Tatiana Tsygankova, Dan Roth
For those languages which use it, capitalization is an important signal for the fundamental NLP tasks of Named Entity Recognition (NER) and Part of Speech (POS) tagging. In fact, i…
Building Low-Resource NER Models Using Non-Speaker Annotation
Tatiana Tsygankova, Francesca Marini, Stephen Mayhew +1
In low-resource natural language processing (NLP), the key problems are a lack of target language training data, and a lack of native speakers to create it. Cross-lingual methods h…
Extending Multilingual BERT to Low-Resource Languages
Zihan Wang, Karthikeyan K, Stephen Mayhew +1
Multilingual BERT (M-BERT) has been a huge success in both supervised and zero-shot cross-lingual transfer learning. However, this success has focused only on the top 104 languages…
Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
Terra Blevins, Stephen Mayhew, Marek Å uppa +11
While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most languages to interrogate these a…
From Tarzan to Tolkien: Controlling the Language Proficiency Level of LLMs for Content Generation
Ali Malik, Stephen Mayhew, Chris Piech +1
We study the problem of controlling the difficulty level of text generated by Large Language Models (LLMs) for contexts where end-users are not fully proficient, such as language l…
Cross-lingual Dataless Classification for Languages with Small Wikipedia Presence
Yangqiu Song, Stephen Mayhew, Dan Roth
This paper presents an approach to classify documents in any language into an English topical label space, without any text categorization training data. The approach, Cross-Lingua…
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark
Stephen Mayhew, Terra Blevins, Shuheng Liu +10
We introduce Universal NER (UNER), an open, community-driven project to develop gold-standard NER benchmarks in many languages. The overarching goal of UNER is to provide high-qual…
On the Strength of Character Language Models for Multilingual Named Entity Recognition
Xiaodong Yu, Stephen Mayhew, Mark Sammons +1
Character-level patterns have been widely used as features in English Named Entity Recognition (NER) systems. However, to date there has been no direct investigation of the inheren…
MasakhaNER: Named Entity Recognition for African Languages
David Ifeoluwa Adelani, Jade Abbott, Graham Neubig +58
We take a step towards addressing the under-representation of the African continent in NLP research by creating the first large publicly available high-quality dataset for named en…