activity
20172025
most citedChoosing Transfer Languages for Cross-Lingual Learning

33 citations · 49 across the 12 of their papers we have counts for

collaborators

17 papers

cs.CL2025

WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects

Daniel Deutsch, Eleftheria Briakou, Isaac Caswell +14

As large language models (LLM) become more and more capable in languages other than English, it is important to collect benchmark datasets in order to evaluate their multilingual p…

cs.CL20223 cited

MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition

David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder +42

African languages are spoken by over a billion people, but are underrepresented in NLP research and development. The challenges impeding progress include the limited availability o…

cs.CL2021

Lexically Aware Semi-Supervised Learning for OCR Post-Correction

Shruti Rijhwani, Daisy Rosenblum, Antonios Anastasopoulos +1

Much of the existing linguistic data in many languages of the world is locked away in non-digitized books and documents. Optical character recognition (OCR) can be used to produce…

cs.CL2021

Dependency Induction Through the Lens of Visual Perception

Ruisi Su, Shruti Rijhwani, Hao Zhu +4

Most previous work on grammar induction focuses on learning phrasal or dependency structure purely from text. However, because the signal provided by text alone is limited, recentl…

cs.CL2021

Evaluating the Morphosyntactic Well-formedness of Generated Texts

Adithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani +4

Text generation systems are ubiquitous in natural language processing applications. However, evaluation of these systems remains a challenge, especially in multilingual settings. I…

cs.CL2021

MasakhaNER: Named Entity Recognition for African Languages

David Ifeoluwa Adelani, Jade Abbott, Graham Neubig +58

We take a step towards addressing the under-representation of the African continent in NLP research by creating the first large publicly available high-quality dataset for named en…