12 papers
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
A. Seza DoÄruöz, Xixian Liao, Verena Blaschke +3
LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human…
An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages
Yinhan Lu, Gaganpreet Jhajj, Chen Zhang +2
In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks from a few examples, making it promising for languages underrepresented in pre-training. Recent…
TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages
Victor Akinode, Senyu Li, Wassim Hamidouche +3
Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, critically underexplored. We in…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation
Idris Abdulmumin, Tajuddeen Gwadabe, Shamsuddeen Hassan Muhammad +11
The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific…
AfrIFact: Cultural Information Retrieval, Evidence Extraction and Fact Checking for African Languages
Israel Abebe Azime, Jesujoba Oluwadara Alabi, Crystina Zhang +16
Assessing the veracity of a claim made online is a complex and important task with real-world implications. When these claims are directed at communities with limited access to inf…