Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…
cs.CL2025
Training language models to be warm and empathetic makes them less reliable and more sycophantic
Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher
Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and compani…
cs.CL2025
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
Franziska Sofia Hafner, Ana Valdivia, Luc Rocher
Language models encode and subsequently perpetuate harmful gendered stereotypes. Research has succeeded in mitigating some of these harms, e.g. by dissociating non-gendered terms s…