11 papers
How Value Induction Reshapes LLM Behaviour
Arnav Arora, Natalie Schluter, Katherine Metcalf +1
Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as h…
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3
Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behaviour is…
Tailored untruths: How personalisation challenges LLM safeguards
João A. Leite, Arnav Arora, Silvia Gargova +5
Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across languages and demographic groups. W…
Revealing Fine-Grained Values and Opinions in Large Language Models
Dustin Wright, Arnav Arora, Nadav Borenstein +3
Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting…
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein
Language embeds information about social, cultural, and political values people hold. Prior work has explored social and potentially harmful biases encoded in Pre-Trained Language…
Investigating Human Values in Online Communities
Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee +1
Studying human values is instrumental for cross-cultural research, enabling a better understanding of preferences and behaviour of society at large and communities therein. To stud…