12 citations · 15 across the 3 of their papers we have counts for
3 papers
MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition
David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder +42
African languages are spoken by over a billion people, but are underrepresented in NLP research and development. The challenges impeding progress include the limited availability o…
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo +6
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval…
Better Than Whitespace: Information Retrieval for Languages without Custom Tokenizers
Odunayo Ogundepo, Xinyu Zhang, Jimmy Lin
Tokenization is a crucial step in information retrieval, especially for lexical matching algorithms, where the quality of indexable tokens directly impacts the effectiveness of a r…