177 citations · 297 across the 48 of their papers we have counts for
44 papers
VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation
Yejin Jeon, Marie Maltais, Virginia Ceccatelli +2
As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization res…
OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages
David Guzmán, Luel Hagos Beyene, Jesujoba Oluwadara Alabi +3
Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially improved synthetic speech quality, yet these gains remain unevenly distributed…
YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset
Peace Busola Falola, Jesujoba O. Alabi, Solomon O. Akinola +3
Named Entity Recognition (NER) is a foundational NLP task, yet research in Yorùbá has been constrained by limited and domain-specific resources. Existing resources, such as Masakha…
Multilinguality as Sense Adaptation
Jan Christian Blaise Cruz, David Ifeoluwa Adelani, Alham Fikri Aji
We approach multilinguality as sense adaptation: aligning latent meaning representations across languages rather than relying solely on shared parameters and scale. In this paper,…
Afri-MCQA: Multimodal Cultural Question Answering for African Languages
Atnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva +16
Africa is home to over one-third of the world's languages, yet remains underrepresented in AI research. We introduce Afri-MCQA, the first Multilingual Cultural Question-Answering b…
Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria's Minority Languages
Oluwadara Kalejaiye, Luel Hagos Beyene, David Ifeoluwa Adelani +4
Nigeria is the most populous country in Africa with a population of more than 200 million people. More than 500 languages are spoken in Nigeria and it is one of the most linguistic…