16 papers
From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa
Mahounan Pericles Adjovi, Victor Olufemi, Roald Eiselen +1
Low-resource African languages lack text corpora needed for language model training. We investigate whether ASR pipelines can extend text resources for two typologically distinct W…
Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability
Mahounan Pericles Adjovi, Roald Eiselen, Prasenjit Mitra
We investigate the translation quality of current large language models (LLMs) for English-to-Hausa and English-to-Fongbe - two typologically distinct West African languages from t…
WAXAL-NET: Finetuned Edge ASR Across 19 African Languages
Victor Tolulope Olufemi, Oreoluwa Babatunde, Ramsey Njema +28
We evaluate whether compact domain-specialized ASR models can outperform massively multilingual foundation models for conversational African speech across 19 languages in the WAXAL…
HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models
Edward Ajayi, Prasenjit Mitra
Evaluating humor in large language models (LLMs) is an open challenge because existing approaches yield isolated, incomparable metrics rather than unified model rankings, making it…
HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation
Edward Ajayi, Prasenjit Mitra
Humor generation poses a significant challenge for Large Language Models (LLMs), because their standard training objective (next-token prediction) inherently conflicts with the sur…
When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP
Mahounan Pericles Adjovi, Roald Eiselen, Prasenjit Mitra
Data scarcity limits NLP development for low-resource African languages. We evaluate two data augmentation methods -- LLM-based generation (Gemini 2.5 Flash) and back-translation (…