activity
20192026
most citedDemographic Inference and Representative Population Estimates from Multilingual Social Media Data

177 citations · 454 across the 73 of their papers we have counts for

collaborators
Showing 2024Show all

16 papers · 1 filter

cs.CL2024

AfriHG: News headline generation for African Languages

Toyib Ogunremi, Serah Akojenu, Anthony Soronnadi +2

This paper introduces AfriHG -- a news headline generation dataset created by combining from XLSum and MasakhaNEWS datasets focusing on 16 languages widely spoken by Africa. We exp…

cs.CL2024

YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text

Akindele Michael Olawole, Jesujoba O. Alabi, Aderonke Busayo Sakpere +1

In this work, we present Yorùbá automatic diacritization (YAD) benchmark dataset for evaluating Yorùbá diacritization systems. In addition, we pre-train text-to-text transformer, T…

cs.CL2024★ 6 cited

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Shivalika Singh, Angelika Romanou, Clémentine Fourrier +21

Cultural biases in multilingual datasets pose significant challenges for their effectiveness as global benchmarks. These biases stem not only from differences in language but also…

cs.CL2024

Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages

Edward Bayes, Israel Abebe Azime, Jesujoba O. Alabi +11

Evaluations of Large Language Models (LLMs) on knowledge-intensive tasks and factual accuracy often focus on high-resource languages primarily because datasets for low-resource lan…

cs.CL2024★ 2 cited

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan +48

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate th…

cs.CL2024

Mitigating Translationese in Low-resource Languages: The Storyboard Approach

Garry Kuwanto, Eno-Abasi E. Urua, Priscilla Amondi Amuok +21

Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which can introduce the translationese effect…