6 papers · 1 filter
Quantifying Media Representation Dynamics Across 25 Years of News Reporting on Policing-related Deaths
Farhan Samir, Jappun Dhillon, Meghna Ravikumar +2
We perform the largest known computational analysis of Canadian news narratives about police-involved deaths, spanning 4,000 articles from the last quarter-century. We develop a no…
When English Rewrites Local Knowledge: Global Narrative Dominance in Large Language Models
Md Arid Hasan, Ruwad Naswan, Farhan Samir +2
Large language models (LLMs) are widely used as cross-lingual knowledge interfaces. However, culturally grounded questions often reflect globally dominant narratives rather than lo…
ZIPA: A family of efficient models for multilingual phone recognition
Jian Zhu, Farhan Samir, Eleanor Chodroff +1
We present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. We first curated IPAPack++, a large-scale…
Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4
Zhuozhuo Joy Liu, Farhan Samir, Mehar Bhatia +2
LLMs have been demonstrated to align with the values of Western or North American cultures. Prior work predominantly showed this effect through leveraging surveys that directly ask…
Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset
Farhan Samir, Emily P. Ahn, Shreya Prakash +3
Curating datasets that span multiple languages is challenging. To make the collection more scalable, researchers often incorporate one or more imperfect classifiers in the process,…
Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on Wikipedia
Farhan Samir, Chan Young Park, Anjalie Field +2
To explain social phenomena and identify systematic biases, much research in computational social science focuses on comparative text analyses. These studies often rely on coarse c…