activity
20242026
collaborators

7 papers

cs.CL2026

AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai +17

Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, ge…

cs.CL2026

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

Naihao Deng, Yilun Zhu, Joan Nwatu +2

Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this w…

cs.CL2026

The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference

Naihao Deng, Alissa Shen, Yiming Feng +5

Large language models (LLMs) are increasingly deployed in multilingual settings, yet the energy costs of serving these models across different languages remain poorly understood. W…

cs.CY2025

Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping

Joan Nwatu, Longju Bai, Oana Ignat +1

Culture shapes the objects people use and for what purposes, yet mainstream Vision-Language (VL) datasets frequently exhibit cultural biases, disproportionately favoring higher-inc…

cs.CY2024

Why AI Is WEIRD and Should Not Be This Way: Towards AI For Everyone, With Everyone, By Everyone

Rada Mihalcea, Oana Ignat, Longju Bai +7

This paper presents a vision for creating AI systems that are inclusive at every stage of development, from data collection to model design and evaluation. We address key limitatio…

cs.CY2024

Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models

Joan Nwatu, Oana Ignat, Rada Mihalcea

Recent work has demonstrated that the unequal representation of cultures and socioeconomic groups in training data leads to biased Large Multi-modal (LMM) models. To improve LMM mo…