11 papers
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
AfrIFact: Cultural Information Retrieval, Evidence Extraction and Fact Checking for African Languages
Israel Abebe Azime, Jesujoba Oluwadara Alabi, Crystina Zhang +16
Assessing the veracity of a claim made online is a complex and important task with real-world implications. When these claims are directed at communities with limited access to inf…
Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages
Badr M. Abdullah, Israel Abebe Azime, Atnafu Lambebo Tonja +14
We present Ethio-ASR, a suite of multilingual CTC-based automatic speech recognition (ASR) models jointly trained on five Ethiopian languages: Amharic, Tigrinya, Oromo, Sidaama, an…
Afri-MCQA: Multimodal Cultural Question Answering for African Languages
Atnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva +16
Africa is home to over one-third of the world's languages, yet remains underrepresented in AI research. We introduce Afri-MCQA, the first Multilingual Cultural Question-Answering b…
Evaluation Sheet for Deep Research: A Use Case for Academic Survey Writing
Israel Abebe Azime, Tadesse Destaw Belay, Atnafu Lambebo Tonja
Large Language Models (LLMs) powered with argentic capabilities are able to do knowledge-intensive tasks without human involvement. A prime example of this tool is Deep research wi…
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation
Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan +32
Translating cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey…