activity
20162022
most citedTaTa: A Multilingual Table-to-Text Dataset for African Languages

3 citations · 3 across the 1 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20223 cited

TaTa: A Multilingual Table-to-Text Dataset for African Languages

Sebastian Gehrmann, Sebastian Ruder, Vitaly Nikolaev +4

Existing data-to-text generation datasets are mostly limited to English. To address this lack of data, we create Table-to-Text in African languages (TaTa), the first large multilin…

cs.CL2021

XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation

Sebastian Ruder, Noah Constant, Jan Botha +8

Machine learning has brought striking advances in multilingual natural language processing capabilities over the past year. For example, the latest techniques have improved the sta…

cs.CL2020

Entity Linking in 100 Languages

Jan A. Botha, Zifei Shan, Daniel Gillick

We propose a new formulation for multilingual entity linking, where language-specific mentions resolve to a language-agnostic Knowledge Base. We train a dual encoder in this new se…

cs.CL2020

Asking without Telling: Exploring Latent Ontologies in Contextual Representations

Julian Michael, Jan A. Botha, Ian Tenney

The success of pretrained contextual encoders, such as ELMo and BERT, has brought a great deal of interest in what these models learn: do they, without explicit supervision, learn…

cs.CL2018

Learning To Split and Rephrase From Wikipedia Edit History

Jan A. Botha, Manaal Faruqui, John Alex +2

Split and rephrase is the task of breaking down a sentence into shorter ones that together convey the same meaning. We extract a rich new dataset for this task by mining Wikipedia'…

cs.CL2017

Natural Language Processing with Small Feed-Forward Networks

Jan A. Botha, Emily Pitler, Ji Ma +5

We show that small and shallow feed-forward neural networks can achieve near state-of-the-art results on a range of unstructured and structured language processing tasks while bein…