5 papers · 1 filter
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…
Framing Migration: A Computational Analysis of UK Parliamentary Discourse
Vahid Ghafouri, Robert McNeil, Teodor Yankov +4
We present a large-scale computational analysis of migration-related discourse in UK parliamentary debates spanning over 75 years and compare it with US congressional discourse. Us…
Into the crossfire: evaluating the use of a language model to crowdsource gun violence reports
Adriano Belisario, Scott A. Hale, Luc Rocher
Gun violence is a pressing human rights issue that affects nearly every dimension of the social fabric, from healthcare and education to psychology and the economy. Reliable data o…
Training language models to be warm and empathetic makes them less reliable and more sycophantic
Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher
Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and compani…
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
Franziska Sofia Hafner, Ana Valdivia, Luc Rocher
Language models encode and subsequently perpetuate harmful gendered stereotypes. Research has succeeded in mitigating some of these harms, e.g. by dissociating non-gendered terms s…