activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens

Hellina Hailu Nigatu, Bethelhem Yemane Mamo, Bontu Fufa Balcha +5

As low-resourced languages are increasingly incorporated into NLP research, there is an emphasis on collecting large-scale datasets. But in prioritizing quantity over quality, we r…

cs.CL2025

A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez Script

Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Henok Biadglign Ademtew +4

Homophone normalization, where characters that have the same sound in a writing script are mapped to one character, is a pre-processing step applied in Amharic Natural Language Pro…

cs.CL2024

A Capabilities Approach to Studying Bias and Harm in Language Technologies

Hellina Hailu Nigatu, Zeerak Talat

Mainstream Natural Language Processing (NLP) research has ignored the majority of the world's languages. In moving from excluding the majority of the world's languages to blindly a…

cs.CL2024

The Zeno's Paradox of `Low-Resource' Languages

Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman +2

The disparity in the languages commonly studied in Natural Language Processing (NLP) is typically reflected by referring to languages as low vs high-resourced. However, there is li…

cs.CL20231 cited

The Less the Merrier? Investigating Language Representation in Multilingual Models

Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Jugal Kalita

Multilingual Language Models offer a way to incorporate multiple languages in one model and utilize cross-language transfer learning to improve performance for different Natural La…