5 papers · 1 filter
Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models
William Guey, Pierrick Bougault, Wei Zhang +2
Emerging AI regulation mandates bias audits of high-risk systems, and audit scores are beginning to be used to rank models. Both uses assume different audit tools measure the same…
Same question, different history: language, national identity, and credit in large language models
William Guey, Pierrick Bougault, Wei Zhang +2
Who invented the radio, Russia's Alexander Popov or Italy's Guglielmo Marconi? Was the telephone the achievement of Bell in the United States or Meucci in Italy? Does printing belo…
Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models
William Guey, Wei Zhang, Pierrick Bougault +4
Large language models (LLMs) are rapidly being integrated into high-stakes public safety systems, including emergency call triage and dispatch decision support, yet their demograph…
BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts
William Guey, Wei Zhang, Pei-Luen Patrick Rau +4
Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hirin…
Mapping Geopolitical Bias in 11 Large Language Models: A Bilingual, Dual-Framing Analysis of U.S.-China Tensions
William Guey, Wei Zhang, Pierrick Bougault +2
Large language models are how hundreds of millions of people now encounter contested political questions, raising a subtle measurement problem: a model that simply agrees with what…