activity
20202026
most citedNL-Augmenter: A Framework for Task-Sensitive Natural Language Augmentation

25 citations · 55 across the 38 of their papers we have counts for

collaborators
Showing 2025 · cs.CLShow all

9 papers · 2 filters

cs.CL2025

Swivuriso: The South African Next Voices Multilingual Speech Dataset

Vukosi Marivate, Kayode Olaleye, Sitwala Mundia +19

This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automa…

cs.CL2025

The Rise of AfricaNLP: A Survey of Contributions, Contributors, Community Impact, and Bibliometric Analysis

Tadesse Destaw Belay, Kedir Yassin Hussen, Sukairaj Hafiz Imam +11

Natural Language Processing (NLP) is undergoing constant transformation, as Large Language Models (LLMs) are driving daily breakthroughs in research and practice. In this regard, t…

cs.CL2025

Mafoko: Structuring and Building Open Multilingual Terminologies for South African NLP

Vukosi Marivate, Isheanesu Dzingirai, Fiskani Banda +9

The critical lack of structured terminological data for South Africa's official languages hampers progress in multilingual NLP, despite the existence of numerous government and aca…

cs.CL2025

HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing

Shamsuddeen Hassan Muhammad, Ibrahim Said Ahmad, Idris Abdulmumin +9

Hausa Natural Language Processing (NLP) has gained increasing attention in recent years, yet remains understudied as a low-resource language despite having over 120 million first-l…

cs.CL2025

HausaNLP at SemEval-2025 Task 2: Entity-Aware Fine-tuning vs. Prompt Engineering in Entity-Aware Machine Translation

Abdulhamid Abubakar, Hamidatu Abdulkadir, Ibrahim Rabiu Abdullahi +9

This paper presents our findings for SemEval 2025 Task 2, a shared task on entity-aware machine translation (EA-MT). The goal of this task is to develop translation models that can…

cs.CL2025

AfroXLMR-Comet: Multilingual Knowledge Distillation with Attention Matching for Low-Resource languages

Joshua Sakthivel Raju, Sanjay S, Jaskaran Singh Walia +2

Language model compression through knowledge distillation has emerged as a promising approach for deploying large language models in resource-constrained environments. However, exi…