activity
20242026
collaborators

36 papers

cs.SD2026

VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

Yejin Jeon, Marie Maltais, Virginia Ceccatelli +2

As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization res…

cs.CL2026

AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages

Hao Yu, Tianyi Xu, Michael A. Hedderich +3

Large language models (LLMs) are increasingly multilingual, yet open models continue to underperform relative to proprietary systems, with the gap most pronounced for African langu…

cs.CL2026

AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages

Happy Buzaaba, Cheikh Mouhamadou Bamba Dione, David Ifeoluwa Adelani +15

Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introdu…

cs.SD2026

NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages

Marie Maltais, Yejin Jeon, Min Ma +7

Speech translation for low-resource languages remains fundamentally limited by the scarcity of high-quality, diverse parallel speech data, a challenge that is especially pronounced…

cs.CL2026

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

David Guzmán, Luel Hagos Beyene, Jesujoba Oluwadara Alabi +3

Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially improved synthetic speech quality, yet these gains remain unevenly distributed…

cs.SD2026

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

Virginia Ceccatelli, Yejin Jeon, David Ifeoluwa Adelani

Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on monolingual, text-based harmful…