activity
20242026
collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2026

How Value Induction Reshapes LLM Behaviour

Arnav Arora, Natalie Schluter, Katherine Metcalf +1

Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as h…

cs.CL2026

BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation

Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3

Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behaviour is…

cs.CL2025

Tailored untruths: How personalisation challenges LLM safeguards

João A. Leite, Arnav Arora, Silvia Gargova +5

Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across languages and demographic groups. W…

cs.CL2025

Revealing Fine-Grained Values and Opinions in Large Language Models

Dustin Wright, Arnav Arora, Nadav Borenstein +3

Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting…

cs.CL2025

Probing Pre-Trained Language Models for Cross-Cultural Differences in Values

Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein

Language embeds information about social, cultural, and political values people hold. Prior work has explored social and potentially harmful biases encoded in Pre-Trained Language…

cs.CL2025

Multi-Modal Framing Analysis of News

Arnav Arora, Srishti Yadav, Maria Antoniak +2

Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its recep…