5 papers
AfrIFact: Cultural Information Retrieval, Evidence Extraction and Fact Checking for African Languages
Israel Abebe Azime, Jesujoba Oluwadara Alabi, Crystina Zhang +16
Assessing the veracity of a claim made online is a complex and important task with real-world implications. When these claims are directed at communities with limited access to inf…
Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages
Badr M. Abdullah, Israel Abebe Azime, Atnafu Lambebo Tonja +14
We present Ethio-ASR, a suite of multilingual CTC-based automatic speech recognition (ASR) models jointly trained on five Ethiopian languages: Amharic, Tigrinya, Oromo, Sidaama, an…
Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens
Hellina Hailu Nigatu, Bethelhem Yemane Mamo, Bontu Fufa Balcha +5
As low-resourced languages are increasingly incorporated into NLP research, there is an emphasis on collecting large-scale datasets. But in prioritizing quantity over quality, we r…
ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding
Israel Abebe Azime, Atnafu Lambebo Tonja, Tadesse Destaw Belay +11
With the rapid development of evaluation datasets to assess LLMs understanding across a wide range of subjects and domains, identifying a suitable language understanding benchmark…
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo +73
Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge pres…