activity
20242026
collaborators

11 papers

cs.CL2026

How Value Induction Reshapes LLM Behaviour

Arnav Arora, Natalie Schluter, Katherine Metcalf +1

Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as h…

cs.CL2026

BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation

Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3

Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behaviour is…

cs.CL2025

Tailored untruths: How personalisation challenges LLM safeguards

João A. Leite, Arnav Arora, Silvia Gargova +5

Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across languages and demographic groups. W…

cs.CL2025

Revealing Fine-Grained Values and Opinions in Large Language Models

Dustin Wright, Arnav Arora, Nadav Borenstein +3

Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting…

cs.CL2025

Probing Pre-Trained Language Models for Cross-Cultural Differences in Values

Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein

Language embeds information about social, cultural, and political values people hold. Prior work has explored social and potentially harmful biases encoded in Pre-Trained Language…

cs.SI2025

Investigating Human Values in Online Communities

Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee +1

Studying human values is instrumental for cross-cultural research, enabling a better understanding of preferences and behaviour of society at large and communities therein. To stud…