papers

Publications (13)

cs.CL2020

Cross-Lingual Ability of Multilingual BERT: An Empirical Study

Karthikeyan K, Zihan Wang, Stephen Mayhew +1

Recent work has exhibited the surprising cross-lingual abilities of multilingual BERT (M-BERT) -- surprising since it is trained without any cross-lingual objective and with no ali…

cs.CL2016

Transliteration in Any Language with Surrogate Languages

Stephen Mayhew, Christos Christodoulopoulos, Dan Roth

We introduce a method for transliteration generation that can produce transliterations in every language. Where previous results are only as multilingual as Wikipedia, we show how…

cs.CL2019

Robust Named Entity Recognition with Truecasing Pretraining

Stephen Mayhew, Nitish Gupta, Dan Roth

Although modern named entity recognition (NER) systems show impressive performance on standard datasets, they perform poorly when presented with noisy data. In particular, capitali…

cs.CL2019

Named Entity Recognition with Partially Annotated Training Data

Stephen Mayhew, Snigdha Chaturvedi, Chen-Tse Tsai +1

Supervised machine learning assumes the availability of fully-labeled data, but in many cases, such as low-resource languages, the only data available is partially annotated. We st…

cs.CL2019

ner and pos when nothing is capitalized

Stephen Mayhew, Tatiana Tsygankova, Dan Roth

For those languages which use it, capitalization is an important signal for the fundamental NLP tasks of Named Entity Recognition (NER) and Part of Speech (POS) tagging. In fact, i…

cs.CL2021

Building Low-Resource NER Models Using Non-Speaker Annotation

Tatiana Tsygankova, Francesca Marini, Stephen Mayhew +1

In low-resource natural language processing (NLP), the key problems are a lack of target language training data, and a lack of native speakers to create it. Cross-lingual methods h…

cs.CL2020

Extending Multilingual BERT to Low-Resource Languages

Zihan Wang, Karthikeyan K, Stephen Mayhew +1

Multilingual BERT (M-BERT) has been a huge success in both supervised and zero-shot cross-lingual transfer learning. However, this success has focused only on the top 104 languages…

cs.CL2026

Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark

Terra Blevins, Stephen Mayhew, Marek Å uppa +11

While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most languages to interrogate these a…

cs.CL2024

From Tarzan to Tolkien: Controlling the Language Proficiency Level of LLMs for Content Generation

Ali Malik, Stephen Mayhew, Chris Piech +1

We study the problem of controlling the difficulty level of text generated by Large Language Models (LLMs) for contexts where end-users are not fully proficient, such as language l…

cs.CL2016

Cross-lingual Dataless Classification for Languages with Small Wikipedia Presence

Yangqiu Song, Stephen Mayhew, Dan Roth

This paper presents an approach to classify documents in any language into an English topical label space, without any text categorization training data. The approach, Cross-Lingua…

cs.CL2024

Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark

Stephen Mayhew, Terra Blevins, Shuheng Liu +10

We introduce Universal NER (UNER), an open, community-driven project to develop gold-standard NER benchmarks in many languages. The overarching goal of UNER is to provide high-qual…

cs.CL2018

On the Strength of Character Language Models for Multilingual Named Entity Recognition

Xiaodong Yu, Stephen Mayhew, Mark Sammons +1

Character-level patterns have been widely used as features in English Named Entity Recognition (NER) systems. However, to date there has been no direct investigation of the inheren…

cs.CL2021

MasakhaNER: Named Entity Recognition for African Languages

David Ifeoluwa Adelani, Jade Abbott, Graham Neubig +58

We take a step towards addressing the under-representation of the African continent in NLP research by creating the first large publicly available high-quality dataset for named en…