collaborators

7 papers

cs.CL2026

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…

cs.CR2026

Computer Science Conferences Should Require Nonrepudiable Experimental Results

Mamadou K. Keita, Christopher Homan

This position paper argues that computer science conferences should require tamper-evident, nonrepudiable attestations of experimental results. We name the underlying problem exper…

cs.LG2026

NSL-MT: Linguistically Informed Negative Samples for Efficient Machine Translation in Low-Resource Languages

Mamadou K. Keita, Christopher Homan, Huy Le

We introduce negative space learning machine translation (NSL-MT), a training method for underresourced languages, that augments limited parallel data with synthetically generated…

cs.CL2026

Where Are We At with Automatic Speech Recognition for the Bambara Language?

Seydou Diallo, Yacouba Diarra, Mamadou K. Keita +3

This paper introduces the first standardized benchmark for evaluating Automatic Speech Recognition (ASR) in the Bambara language, utilizing one hour of professionally recorded Mali…

cs.LG2026

InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages

Mamadou K. Keita, Sebastien Diarra, Christopher Homan +1

Effective text generation and chat interfaces for low-resource languages (LRLs) remain a challenge for state-of-the-art large language models (LLMs) to support. This is mainly due…

cs.CL2026

Grammatical Error Correction for Low-Resource Languages: The Case of Zarma

Mamadou K. Keita, Adwoa Bremang, Huy Le +3

Grammatical error correction (GEC) aims to improve text quality and readability. Previous work on the task focused primarily on high-resource languages, while low-resource language…