4 citations · 7 across the 5 of their papers we have counts for
5 papers
Leveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource Languages
Idris Abdulmumin, Auwal Abubakar Khalid, Shamsuddeen Hassan Muhammad +5
The importance of qualitative parallel data in machine translation has long been determined but it has always been very difficult to obtain such in sufficient quantity for the majo…
HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language
Shantipriya Parida, Idris Abdulmumin, Shamsuddeen Hassan Muhammad +7
This paper presents HaVQA, the first multimodal dataset for visual question-answering (VQA) tasks in the Hausa language. The dataset was created by manually translating 6,022 Engli…
SemEval-2023 Task 12: Sentiment Analysis for African Languages (AfriSenti-SemEval)
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Seid Muhie Yimam +7
We present the first Africentric SemEval Shared task, Sentiment Analysis for African Languages (AfriSenti-SemEval) - The dataset is available at https://github.com/afrisenti-semeva…
HausaNLP at SemEval-2023 Task 10: Transfer Learning, Synthetic Data and Side-Information for Multi-Level Sexism Classification
Saminu Mohammad Aliyu, Idris Abdulmumin, Shamsuddeen Hassan Muhammad +4
We present the findings of our participation in the SemEval-2023 Task 10: Explainable Detection of Online Sexism (EDOS) task, a shared task on offensive language (sexism) detection…
The African Stopwords project: curating stopwords for African languages
Chris Emezue, Hellina Nigatu, Cynthia Thinwa +19
Stopwords are fundamental in Natural Language Processing (NLP) techniques for information retrieval. One of the common tasks in preprocessing of text data is the removal of stopwor…